
H Company Holo4: 27B Computer-Use Agent Hits 61.7% on OSWorld 2.0
H Company launched the Holo4 series of computer-use agents in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts, available immediately via the H Models API with weights on Hugging Face. Holo4-27B scores 61.7% on OSWorld 2.0 versus 81.8% for Opus 5.5, at a fraction of the cost per task. The models support GUI, code, MCP, and API interfaces across desktop, web, and Android.
Executive Overview
H Company has officially announced the launch of the Holo4 series of agentic models, available today via the H Models API with weights published on Hugging Face. The series ships in two sizes: Holo4-27B (dense) and Holo4-35B-A3B (Mixture of Experts), alongside an updated Holotron4 Nano. Holo4-27B scores 61.7% on OSWorld 2.0 — the hardest desktop-control benchmark — versus 81.8% for Opus 5.5, while H Company claims a dramatically lower cost per task. The models interact with software through any available interface: GUIs, code, MCP, and APIs, automatically selecting the best interface for each task.
📊 Official Data & Technical Specifications Sheet
| Technical Axis | Confirmed Official Data |
|---|---|
| 💰 Pricing & Usage Cost | Holo4 is priced via the H Models API with official token pricing (single run). The company confirms cost per task is orders of magnitude lower than frontier models on OSWorld 2.0 and AutomationBench. For comparison: Qwen3.8 27B is priced at Alibaba Cloud rates with a 20% discount on input cache hits. |
| 🌐 Platforms & Immediate Availability | Available immediately on the H Models API. Weights on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF formats. Runs on: desktop, web, Android, isolated code sandbox, and business APIs. |
| ⚡ Performance & Speed Benchmarks | OSWorld 2.0: Holo4-27B = 61.7% | Holo4-35B-A3B = 30.9% | Opus 5.5 = 81.8% | Opus 5 = 70.2% | GPT-5.6 Sol = 66.2%. AutomationBench v1.0.6: internal measurements for Holo4, Qwen3.8 27B, and Qwen3.6 35B-A3B. |
| 🛡️ Security & Breach Resistance | No specific numeric security rating was stated in the official announcement. The company notes the harness was rebuilt with trusted memory and a shell on the desktop machine, and trained via reinforcement learning (RL) on environments from the Agentic Task Factory. |
| 🧠 Context Window | The context window size was not disclosed numerically in the announcement. The company confirms the new harness manages context over hundreds of steps with trusted memory. |
| 🌍 Arabic Language & Regional Support | No explicit Arabic support was mentioned in the official announcement. The models are built on the Qwen base (Qwen3.8 27B and Qwen3.6 35B-A3B), known for multilingual support, with global availability via the H Models API and Hugging Face. |
Deep-Dive Features & Architecture
Holo4 is built on H Company's previous model and interacts with software through any available interface: graphical user interfaces (GUIs), code, the MCP protocol, and application programming interfaces (APIs). The model automatically selects the most suitable interface for the task — unlike most agentic models, which are trained for a single interface only. Holo4 was trained via supervised learning and reinforcement learning across a wide range of environments and tasks, including tasks generated by the company's own Agentic Task Factory, which has so far produced approximately 10,000 tasks across web applications, MCP servers, and desktop environments — including hybrid environments that present the same state through both GUI and MCP simultaneously.
The company rebuilt its harness — the loop that executes the model's actions and manages its context over hundreds of steps — based on agentic performance observations on OSWorld 2.0. Key changes include granting the agent trusted memory that tracks hundreds of steps and providing a shell on the desktop machine itself. The same post-training technique was applied to the Nemotron 3 Nano Omni model as part of H Company's membership in the NVIDIA Nemotron Coalition, producing Holotron4 Nano, which achieves absolute percentage-point improvements on GUI tasks and MCP, API, and code environments.
Benchmark & Competitive Performance
On the OSWorld 2.0 benchmark — the hardest for desktop control — Holo4-27B scores 61.7% versus 81.8% for Opus 5.5, 70.2% for Opus 5, and 66.2% for GPT-5.6 Sol. Holo4-35B-A3B scores 30.9%. H Company confirms these results are achieved with orders of magnitude fewer parameters and a far lower cost per task. On AutomationBench v1.0.6, internal measurements cover Holo4, Qwen3.8 27B, and Qwen3.6 35B-A3B.
Industry Impact & Enterprise Adoption
The Holo4 launch signals a shift toward multi-interface computer-use agents that can operate across desktop, web, Android, code sandboxes, and business APIs without requiring separate models per platform. With open weights on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF formats, enterprises can self-host or deploy via the H Models API. The cost-efficiency claim — orders of magnitude lower cost per task than frontier models — positions Holo4 as a practical option for large-scale desktop automation and agentic workflows, even where raw benchmark scores trail the largest proprietary models.
Conclusion
Holo4 represents a focused bet on cost-efficient, multi-interface computer-use agents. While Holo4-27B's 61.7% on OSWorld 2.0 trails Opus 5.5's 81.8%, the combination of open weights, immediate API availability, and dramatically lower per-task cost makes it a compelling option for enterprises scaling agentic automation. The updated Holotron4 Nano extends the same post-training stack to smaller deployments, broadening the addressable use cases.
Media Source: Hugging Face | البيان الرسمي للشركة: المصدر الأصلي | Fact Verification & Analysis: AI Tools Oasis
Frequently Asked Questions
Holo4 is a new series of agentic models developed by H Company, available in two sizes: a 27B dense model and a 35B-A3B Mixture of Experts, plus an updated Holotron4 Nano. It interacts with software through GUI, code, MCP, and API interfaces.
Holo4-27B scores 61.7% on OSWorld 2.0, compared to 81.8% for Opus 5.5, 66.2% for GPT-5.6 Sol, and 70.2% for Opus 5. Holo4-35B-A3B scores 30.9% on the same benchmark.
Holo4 is available immediately via the H Models API, with weights published on Hugging Face in BF16, FP8, NVFP4, and 4-bit GGUF formats, alongside the smaller Holotron4 Nano. Execution trajectories are also available for review and download.
Holo4 runs on desktops, the web, Android, isolated code sandboxes, and business APIs. The same model is invoked the same way without needing to select a different model for each platform.
Holotron4 Nano is an updated version of Holotron 3, built with the same post-training stack used for Holo4 on the Nemotron 3 Nano Omni model under the NVIDIA Nemotron Coalition, delivering improvements on GUI tasks and MCP, API, and code environments.

AI Tools Oasis Team
Bringing you the latest news and analysis in the world of Artificial Intelligence with accuracy and credibility. Follow us for all updates.
