
HRM-Text
Open-sourced in May 2026, HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.
Key Traits
Data-Efficient Training
Trained on ~40B tokens, using up to 1000× less data than the 4–36T tokens used by the models we benchmark against.
Compact Yet Powerful
Built with 1.15B parameters while remaining competitive with models several times its size on reasoning-heavy benchmarks.
Native Edge Reasoning
Runs locally with a 0.6 GiB footprint at int4 quantization, enabling advanced reasoning without cloud dependency.
Application Domains





Application Domains
Our architecture powers advanced reasoning across complex, high-impact real-world domains.
Benchmarks
HRM-Text is a proof-of-concept model with no post-training. The numbers below reflect architecture performance alone.
Tokens vs benchmark average
Upper-left is better: fewer training tokens, higher benchmark averageFLOPs vs benchmark average
Upper-left is better: lower training FLOPs, higher benchmark average HRM FLOPs include the H-L recurrent forwards.Despite its compact size, HRM-Text delivers competitive results across reasoning benchmarks, including 56.2% on MATH, 81.9% on ARC-Challenge, 82.2% on DROP, and 60.7% on MMLU.
Benchmark Breakdown
A benchmark that tests mathematical reasoning and problem solving, often requiring multi-step logic rather than simple recall.
A reading comprehension benchmark that tests a model’s ability to reason over passages, especially with numbers, counting, comparison, and discrete operations.
The AI2 Reasoning Challenge- Challenge Set, designed to test science reasoning through difficult grade-school science questions that require inference and commonsense understanding.
Massive Multitask Language Understanding, a broad benchmark covering many subjects, used to evaluate general knowledge and multi-domain reasoning.
