Holo4: powering generalist computer-use agents
- ID
- 29251
- Status
- summarized
- Published
- 28 Sep 2026, 5:44 PM
- Fetched
- 28 Sep 2026, 7:30 PM
- Provider
- Hugging Face Blog
- Category
- developer-ai
- Original URL
- https://huggingface.co/blog/Hcompany/holo4
- Source URL
- https://huggingface.co/blog/feed.xml
Summary
- Score
- 6.5
- Created
- 28 Sep 2026, 7:30 PM
- Tags
- Audience
- developersai_ml_learnersai_agent_usersvibe_coderssaas_founders
What happened
H Company released Holo4, a family of computer-use agent models in two sizes — 27B dense and 35B-A3B Mixture of Experts — plus Holotron4 Nano, an updated Holotron 3, all served on the H Models API with FP16, FP8 and GGUF weights on Hugging Face. The same model drives GUIs, writes and runs its own code, and calls MCP or API tools rather than needing a separate model per interface, trained via supervised and reinforcement learning on environments including ones generated by their Agentic Task Factory. On OSWorld 2.0 the 27B scores 61.7% against 81.8% for Opus 5.5, while the larger 35B-A3B MoE reaches only 30.9%, and every trajectory behind the published scores is open-sourced for replay or download.
Why it matters
The open weights plus open trajectory dataset mean you can self-host a computer-use agent or fine-tune on their published steps instead of paying frontier API rates — but the size naming is a trap: the 35B-A3B MoE scores 30.9% on OSWorld 2.0 versus 61.7% for the 27B dense, so defaulting to the 'bigger' model for GUI work costs you roughly half the success rate. Pick the 27B dense or Holotron4 Nano for screen-based tasks, and check the FP8/GGUF builds against your own workflow before committing.
Discussion angle
Why does a 35B-A3B MoE land at 30.9% on OSWorld 2.0 while a 27B dense model hits 61.7% — is active-parameter budget the bottleneck for GUI agents, and does the published trajectory dataset give a small team enough to fine-tune a local model that closes the gap to Opus 5.5's 81.8%?