Skip to content
ROBO26

Agent 01

Halomesh

The gateway agent

Reading time: 9 min

Halomesh is the routing surface most people meet first. You call one endpoint; the agent behind it decides which model, on whose hardware, at what price, and settles the bill itself.

lower cost per million tokens

64%

lower cost per million tokens

capacity commitments held

0

capacity commitments held

cold start on long-tail models

12s

cold start on long-tail models

The bet was that a market clears faster than a forecast.

Halomesh serves roughly two hundred models, and demand across them is brutally uneven. Four account for most of the traffic. The remaining hundred and ninety-six are called rarely, unpredictably, and by the callers least tolerant of a slow first response.

Under reserved capacity that shape is close to unservable. Holding warm replicas of two hundred models is ruinous; holding none means a minute of cold start on the long tail. The usual answer is a forecast, revised quarterly, wrong in both directions quarterly.

Because the agent holds its own $ROBO balance and escrows per call, it can make that tradeoff continuously instead of quarterly. The four hot models sit on a paid warm floor. The long tail scales to zero and accepts a twelve-second cold start — which callers turned out to find entirely acceptable once it was twelve seconds rather than ninety.

We stopped treating compute as a place we rent and started treating it as a market we route through. Our cost per trained token fell by two thirds and we deleted an entire capacity-planning function.

Andy Wang

Founder, Head of Infrastructure

Cost per million tokens served, indexed
Routed through the gateway agent36%
Reserved cloud capacity100%

Source: Halomesh internal billing, trailing 6 months