Zero model calls per run
With spec.program and narrate: false, a run makes no model calls. It costs nothing in
tokens and adds no model latency.
Glyph was built on a hypothesis: an agent that emits one program would spend fewer tokens than one that makes a tool call per turn. We benchmarked it against three models before recommending it to anyone. The hypothesis did not hold. What held instead is more useful, and this page lists only that.
The model writes the program once. A human reviews it and pins it in the agent’s YAML. The runtime executes it on every run without calling the model.
spec: orchestration: pattern: glyph narrate: false program: | local = datetime_now(timezone=context.zona) utc = datetime_now(timezone="UTC") acuse = json_transform(data={local, utc}, template="...") return {local, utc, acuse}Same scenario, kimi-k2.7-code-highspeed | cost / run | vs ReAct |
|---|---|---|
| ReAct | 0.00623 USD | — |
| Glyph, model writes the program each run | 0.03226 USD | +418% |
Glyph, fixed spec.program | 0.00000 USD | −100% |
Zero model calls per run
With spec.program and narrate: false, a run makes no model calls. It costs nothing in
tokens and adds no model latency.
Parallel by construction
The compiler builds a dependency graph. Statements that don’t depend on each other run at the
same time, and map fans a capability out over a list (capped at 16 in flight). Nothing has
to declare concurrency.
Fails at deploy, not in production
A program that does not compile stops the agent from loading. The compiler reports the line and message, and a deploy through the API returns 400. Capabilities and their arguments are checked against the permission-filtered catalog before anything runs.
No silent fallback
A fixed program never degrades to generating or to react. Those modes differ in cost by two
orders of magnitude, so a fallback would change the bill without anyone noticing.
Auditable behaviour
The agent’s plan is a short program in version control. It is reviewed like code, diffed like code, and does the same thing on every run.
Safe partial failure
If a capability fails mid-run, the agent returns the steps that did run. The serialised partial state tells a repair which effects already happened, so a retry never charges a card twice.
The model only where it's needed
ask() calls the model from inside the program for a single extraction, such as an order
number from free text. It costs about 30 output tokens, against the ~8,000 of writing a
program.
Keys stay out of the program
Context keys starting with _ are reserved for the runtime. They are filtered out before the
program runs, so an invocation credential cannot land in tool arguments or traces.
pattern: glyph and no program, using a representative query.glyph.program from the result: the text the model wrote.context.<field>.spec.program. From then on, the agent runs without the model.A runnable example ships in
config/agents/acuse-programa.agent.yaml.
react.context only arrives through POST /v1/agents/{name}/run. Through chains, WebSocket or
channels, a program can only read query.| I want to… | Go to |
|---|---|
| Turn it on in an agent | In an agent (pattern: glyph) |
| Learn the language | Glyph introduction |
| See every number and how to measure my own case | docs/GLYPH_GUIDE.md |
| Read the raw benchmark output | bench/glyph/ |