Unreal Labs has released Unreal Agent, an open-source agent harness designed to reduce the cost and latency overhead created by tool calls in AI agent workflows. Released on September 22, the Go-based project uses asynchronous tool execution so an agent can continue working while long-running operations run in the background.
Unreal Labs says its current implementation can reduce costs by up to 40% compared with Codex and up to 20% compared with Pi across its production workloads and agentic benchmarks. Those figures come from Unreal Labs’ own testing rather than an independent evaluation, but the company has published the benchmark configurations, costs and run identifiers needed for others to reproduce the experiments.
The release highlights a part of AI agent infrastructure that is often less visible than the underlying model: the harness that manages the model’s interaction with tools.
Quick Summary
- Unreal Labs launched Unreal Agent, an open-source AI agent harness.
- The system is designed around asynchronous tool execution.
- Long-running tools can continue working while the agent performs other tasks.
- Unreal Labs reports lower benchmark costs compared with its Codex and Pi comparisons.
- On Terminal-Bench 4.0, Unreal Labs reports the same 57.9% pass rate as its Codex comparison at a lower reported cost.
- The project is open source and MIT licensed.
- Unreal Agent is a harness, not a new AI model.
- Its main significance is the potential to reduce tool-call and model-turn overhead in agentic workflows.
What Is Unreal Agent?
Unreal Agent is not a new large language model. It is an agent harness that sits between an AI model and the tools it uses to complete tasks.
The system is designed around asynchronous execution. Instead of requiring the model to wait for every tool operation to finish before continuing, Unreal Agent records a tool operation as in progress and allows execution to continue in the background. When the operation finishes, its result is added to the session history and another model call can process the result.
That architecture is intended to address a common source of overhead in AI agents. Long-running operations such as development-environment setup, builds, tests or other shell commands can otherwise force an agent to spend additional turns monitoring whether work has completed.
Unreal Labs says its design allows users to steer an agent without having to wait for outstanding tool calls to finish. It also allows the agent to schedule additional useful tool work between model calls.
How Unreal Agent’s Async Architecture Works?
The central change is the separation between tool-call coordination and tool execution.
When a model requests a tool, Unreal Agent can immediately record the operation as in progress while the actual work is handled separately. Once the operation completes, its result becomes available to the model through the session history.
The project’s architecture includes a coordinator, session store, context builder, LLM adapter, tool registry, tool translators and an operation manager. The repository describes operations as serializable units of work that can be dispatched to an execution system independently of the coordinator’s model loop.
The approach also keeps the harness relatively small. Unreal Labs says its cost-efficiency strategy relies partly on simple prompts, token-efficient tool results and avoiding sub-agents or additional workflows. The company also emphasizes doing more tool work per model turn rather than spending model tokens on polling and waiting.
The current project includes fixed tool definitions for Bash, ViewImage and skill use, while its architecture allows alternative implementations of several underlying components.
Unreal Agent Benchmark Results
Unreal Labs tested the harness using GPT-6 Astra at xhigh reasoning effort and compared it with Codex and Pi across several agent benchmarks. The company reports broadly similar or somewhat higher task performance while using fewer resources in its tests.
| Benchmark | Unreal Agent | Codex | Pi |
|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% / $1,428 | 57.9% / $2,350 | 55.0% / $1,827 |
| SWE-Atlas Codebase QnA | 65.8% / $936 | 63.3% / $1,303 | 64.0% / $1,033 |
| DeepSWE 1.1 | 72.4% / $1,367 | 69.0% / $1,633 | 69.6% / $1,584 |
| Agents’ Last Exam — ALE-CLI | 30.0% / $217 | 29.0% / $292 | 29.0% / $262 |
The most notable result is Terminal-Bench 4.0, where Unreal Agent and the Codex leaderboard baseline both recorded a 57.9% pass rate in the company’s comparison. Unreal Labs reports a total benchmark cost of $1,428 for Unreal Agent compared with $2,350 for the Codex baseline, a difference of roughly 39%.
Terminal-Bench 4.0 itself is a continuously maintained benchmark for evaluating AI agents on complex terminal tasks. Its latest version changed task resources and removed saturated tasks, meaning results should be interpreted within the specific 4.0 evaluation environment rather than treated as a universal measure of coding-agent performance.
Unreal Labs also reported lower costs on SWE-Atlas Codebase QnA and DeepSWE 1.1 while recording higher pass rates than its Codex comparison runs. On Agents’ Last Exam, the full-pass rate was only marginally different from the other systems, while the reported cost was lower. The company notes that marginal pass-rate differences can result from benchmark variance.
The 40% Cost-Saving Claim Needs Context
The headline 40% cost reduction should not be interpreted as proof that every AI agent will become 40% cheaper by adopting Unreal Agent.
The figure comes from Unreal Labs’ own measurements. The comparison also involves particular model settings, workloads and benchmark configurations. The company’s Terminal-Bench comparison, for example, uses GPT-6 Astra at xhigh reasoning effort, while the Codex figure is identified as the leaderboard baseline.
That distinction matters because agent costs depend on much more than the harness itself. Model selection, reasoning effort, prompt size, tool usage, output length, task difficulty and execution environment can all affect total spending.
The more defensible takeaway is that agent architecture can materially affect the number of model turns and tokens required to complete a task. Unreal Agent provides one open implementation of that idea, with public benchmark runs that developers can inspect and reproduce.
Unreal Agent Is Fully Open Source
Unreal Agent is available publicly on GitHub as a Go project. The Unreal Labs repository describes it as an async-first agent harness and includes the harness library, command-line executables and benchmark infrastructure.
The Unreal Labs GitHub organization lists the repository under an MIT license, making the project available for developers to inspect, modify and integrate under the terms of that license.
The current SDK provides three main entry points:
- A Go library for integrating the harness into applications
- A runner executable for running agents from the command line
- A Harbor-compatible benchmark runner for evaluating agent performance
This makes the release relevant not only to users looking for a ready-made coding agent but also to developers building their own agent infrastructure.
Why Async Execution Matters for AI Agents?
As AI agents move beyond short question-and-answer interactions, the amount of time spent interacting with external tools becomes increasingly important.
A coding agent might need to inspect a repository, install dependencies, start a development server, run tests, execute a build and analyze the resulting logs. Some of these operations can happen independently or take significantly longer than a normal model inference.
A synchronous design can leave the model repeatedly checking whether an operation has completed. An asynchronous harness can instead allow those operations to proceed while the agent performs other useful work.
Unreal Labs’ release therefore focuses on a systems-level optimization rather than introducing another frontier model. The company argues that harness design itself is an important area of AI-agent research, particularly as developers try to reduce the cost of running increasingly complex autonomous workflows.
There are also implementation limitations. Unreal Labs notes that its asynchronous tool-result approach exposed compatibility issues with some inference providers because the relevant tool-call result behavior is not consistently specified across APIs. That means the architecture may require additional engineering when deployed across different model providers.
For developers building AI agents, the release offers a concrete open-source implementation of asynchronous tool orchestration rather than another closed agent product. Its reported benchmark results will need broader independent reproduction, but the underlying design provides a useful example of how agent infrastructure can influence both responsiveness and model usage.
Unreal Agent is therefore notable less for changing the underlying AI model than for changing how an AI agent manages the work around that model.
Also Read –
Kimi Code Desktop: AI Coding Agent for Windows & Mac
Source
Unreal Labs – Unreal Agent announcement
Unreal Agent GitHub repository
Terminal-Bench 4.0 documentation and results


