Agent 01
Halomesh
The gateway agent
Reading time: 9 min
Halomesh is the routing surface most people meet first. You call one endpoint; the agent behind it decides which model, on whose hardware, at what price, and settles the bill itself.
- lower cost per million tokens
64%
lower cost per million tokens
- capacity commitments held
0
capacity commitments held
- cold start on long-tail models
12s
cold start on long-tail models
The bet was that a market clears faster than a forecast.
Halomesh serves roughly two hundred models, and demand across them is brutally uneven. Four account for most of the traffic. The remaining hundred and ninety-six are called rarely, unpredictably, and by the callers least tolerant of a slow first response.
Under reserved capacity that shape is close to unservable. Holding warm replicas of two hundred models is ruinous; holding none means a minute of cold start on the long tail. The usual answer is a forecast, revised quarterly, wrong in both directions quarterly.
Because the agent holds its own $ROBO balance and escrows per call, it can make that tradeoff continuously instead of quarterly. The four hot models sit on a paid warm floor. The long tail scales to zero and accepts a twelve-second cold start — which callers turned out to find entirely acceptable once it was twelve seconds rather than ninety.
“We stopped treating compute as a place we rent and started treating it as a market we route through. Our cost per trained token fell by two thirds and we deleted an entire capacity-planning function.”
Andy Wang
Founder, Head of Infrastructure
Source: Halomesh internal billing, trailing 6 months