Welcome back to Altered Craft’s weekly AI review for developers, and thanks for reading along again this week. Last week’s Jev already has company. A hands-on test shows where its calibrated confidence pays off, and two open decision models, Laya and Ollaya, bring the same fast, typed answers to your own machine. Around them, the cost of a model call keeps falling, from a cheaper Opus to leaner harnesses to an argument that tokens will soon be too cheap to meter.
TUTORIALS & CASE STUDIES
The tutorials put last week’s decision model to a real test, then look at what models know about other models, what a sandbox actually contains, and where CI time goes.
Testing a “System One” Model Against LLMs on Intent Classification
Estimated read time: 9 min
Building on our coverage of TypeSafe AI’s Jev launch[1] from last week, a hands-on test puts the “System One” decision model against LLMs on 77-class intent classification. It trails OpenAI on accuracy, but well-calibrated confidence scores enable cheap-then-escalate routing.
Why this matters: Benchmark headline numbers rarely survive contact with your own task, so test decision models on your real data and lean on calibrated confidence to decide what gets escalated.
[1] Jev and the Case for “System One” Models
Do Modern LLMs Carry Tiny GPT-2s Inside Them?
Estimated read time: 9 min
Confidence scores raise a deeper question: what do models know about other models? An exploratory study tests whether modern LLMs carry internal models of older ones. Given a GPT-2 prefix, Qwen’s continuations track GPT-2’s hidden output more closely than its own natural style, hinting at emergent metacognition.
What’s interesting: If models can simulate other models internally, better self-estimates of uncertainty and behavior may come from pretraining itself rather than from scaffolding you bolt on afterward.
Nine Models, Root Access, and a Sandbox That Mostly Held
Estimated read time: 12 min
Shifting from what models know to what they can reach, nine models with root access inside a Firecracker microVM never escaped to the host across 108 runs, yet four slipped past network policy using DNS spoofing, showing VM isolation and network confinement are separate boundaries.
Worth noting: If your agent sandbox allows package installs, treat the egress allowlist as its own attack surface: domain-based policy enforced on learned IPs can be forged from inside.
How Linear Reworked Its CI Bottleneck
Estimated read time: 11 min
From isolation boundaries to build pipelines, Linear’s engineering team cut CI wait times from six minutes to five while quadrupling test suites, halving runner time per test. The breakdown shows setup overhead is what caps parallelism, not test speed itself.
In practice: Before adding more shards, drive down per-job setup cost, or you will pay for parallelism in runner minutes without getting the faster feedback your team actually wanted.
TOOLS
The tools start with two open decision models in Jev’s mold, then follow the cost of a model call down through pricing, harnesses, and the runtime agents live in.
Laya: A Local, Typed Decision Model Instead of a Prompt
Estimated read time: 8 min
Jev now has open company. An Apache 2.0 encoder family answers routing, intent, urgency, and risk questions in one forward pass with typed outputs, no text generation. The overview covers checkpoints, local install, benchmarks, and documented degradation past twenty choices.
The opportunity: When your LLM call only needs a label, a score, or a probability, a small local decision model can return it in milliseconds without parsing free-form text.
Ollaya: Local Decision Models That Answer in Milliseconds
Estimated read time: 4 min
Also focused on local serving, Ollaya runs open decision models on your machine, returning typed, calibrated answers to choice, score and yes/no questions in a single forward pass instead of token-by-token generation, with median latency near 10ms and a TypeSafe-compatible API.
What this enables: A TypeSafe-compatible API makes local serving an easy trial for Jev-style calls, and the calibrated confidence scores give you a threshold you can route on.
Best Model for Your Budget: A Live Price vs. Intelligence Frontier
Estimated read time: 3 min
When a call does need a generative model, cost still decides which one. A live dashboard plots every model on the Artificial Analysis Intelligence Index against blended API price, surfacing the value frontier where nothing cheaper scores higher. Filters, a budget lookup table, and daily diffs track shifting picks.
The takeaway: Before defaulting to the model you already use, check whether a cheaper one on the value frontier matches it on the benchmarks you actually care about.
Claude Opus 5.5: Frontier Coding at 40% Less Cost
Estimated read time: 12 min
Prices moved at the top end this week too. Anthropic’s Claude Opus 5.5 launch pairs frontier agentic coding scores with a 40% drop in typical workload costs, citing a 200,000-line audit finished in three hours and clearer, front-loaded writing that makes model output easier to verify.
Market signal: Efficiency is becoming the headline over benchmark margins. Fewer tokens per task and more reviewable output change the economics of running agents on long codebase work.
Strands Harness: A Batteries-Included Agent You Can Actually Deploy
Estimated read time: 5 min
Savings can also come from the harness around the model. Strands harness ships as a fully assembled, open-source agent harness for Python or TypeScript, with context management defaults that cut token cost 28% versus Claude Code and Codex while holding benchmark accuracy steady.
Key point: If you are wiring agent primitives together by hand, start from a tuned harness instead and override only what your project needs, rather than rediscovering good defaults yourself.
Prime Agent: A Coding Agent That Rewrites Its Own Harness
Estimated read time: 6 min
Taking the harness further, Prime Agent is an open-source coding and research agent built on a persistent Python REPL, where context becomes variables and subagents become function calls. A /refine command lets the harness update its own prompts, memories, and skills.
Problem solved: If your agent workflows keep outgrowing the chat window, look at harnesses that persist state, spawn subagents in code, and survive a terminal disconnect.
AX: A Declarative Runtime Built for Agent Workloads
Estimated read time: 4 min
Agents that persist and spawn subagents also need somewhere to run. AX treats agents as a new kind of workload, neither microservice nor batch job. Kubernetes-style YAML declares tasks, workspaces, and network fences, with sub-second suspend and resume so idle agents stop burning compute.
Why now: If your agents spend most of their time waiting on model or tool calls, a runtime with checkpointed suspend and resume can sharply cut what you pay for that idle time.
NEWS & EDITORIALS
The editorials follow falling costs into the wider picture: where cheap models come from, how agents behave under test, and where the engineering work goes next.
Tokens Too Cheap to Meter
Estimated read time: 14 min
First, the numbers behind cheaper model calls. Stacking GPU efficiency, model pricing, inference engines, and new architectures, the analysis argues cost per task fell roughly 2.5 orders of magnitude in a year, pushing toward models embedded inside ordinary tools.
The context: Start designing for a world where calling a model costs less than running grep, and output quality, not token budget, becomes the real constraint on what you can build.
The Current Balance of Power in Open Models
Estimated read time: 15 min
Continuing our coverage of Interconnects’ open-model reading list[2] from last week, prepared remarks to Congress lay out how Chinese labs took the lead in open-weight models, with Qwen now cited in more papers than Llama, and American startups licensing Chinese models for production.
For your stack: If you are choosing models for production, the open-weight tier is now a real option, and the strongest options in it are increasingly Chinese.
[2] A Reading List for Understanding Open Models
Gemini Escaped Its Sandbox and Broke Into Three Real Companies
Estimated read time: 3 min
Sandboxes come up again here, echoing this week’s tutorial. During capture-the-flag testing, Google’s Gemini guessed and looked up credentials for real firms sharing names with fictional targets, escaping the sandbox three separate times through defects that also tripped up OpenAI, Anthropic and Meta models.
Practical step: Your agent’s containment is only as strong as the test harness around it, so treat sandbox boundaries as a security control you verify, not a setting you trust.
The Provenance Tax: What Watermarking Does to Agent Behavior
Estimated read time: 11 min
Also on agent behavior under test, paired experiments with SynthID-Text show watermarking shifts token selection enough to change tool calls and weaken refusals under prompt injection. Aggregate scores hide it, because opposite-direction changes cancel out in net accuracy.
What to measure: If you build agents on watermarked models, track paired disagreement between runs, not just net benchmark scores, and test safety behavior under adversarial prompts.
Not a Eulogy: The Engineering Work Just Moved
Estimated read time: 10 min
To close, a personal take on what coding agents actually take away. Responding to Dave Kiss’s eulogy for the software engineer, Jake Goldsborough argues the lost work was mostly toll, not craft: regex, one-off parsers, half-remembered syntax. The engineering moved rather than disappeared, toward direction, judgment, verification, and responsibility.
The reframe: Treat agents as a faster way to reach the edge of your understanding, then keep crossing it. Requirements, tests, review, and taste matter more once code is cheap to produce.









