Glad to have you here for Altered Craft’s weekly AI review for developers, and thank you for reading. For the third week running, decision models lead the edition: OpenAI previewed a Decisions API, and OpenJev put an open-weights version on a single GPU. Their first job is routing, and routing shows up throughout, from a router that cut coding agent costs 64% to a skills library with its own routing table. The editorials ask what follows: once a model picks the route, who stays accountable?
TUTORIALS & CASE STUDIES
The tutorials start with OpenAI’s take on the decision model, then put routing to work inside a real coding agent and cover how to measure and prompt the models you route to.
OpenAI’s Decisions API: When You Need an Answer, Not a Paragraph
Estimated read time: 12 min
Building on our coverage of the hands-on Jev intent-classification test[1] from last week, OpenAI now has its own entry. The limited-preview Decisions API returns a bounded choice instead of prose, turning routing, triage, and agent next-step selection into a typed call. The guide argues model judgment should never replace deterministic policy in production agents.
Key point: Define the answer space before inference, then let code own authority. A decision model can judge what a request means, but your policy layer still grants the permission.
[1] Testing a “System One” Model Against LLMs on Intent Classification
Model Routing Cut Coding Agent Costs by 64%
Estimated read time: 11 min
Putting routing to work, this walkthrough builds a model router inside LangChain’s Open SWE coding agent, showing that routing belongs in the harness, not a generic gateway. An A/B test across 973 threads cut median cost 64% with no quality drop.
In practice: Classify incoming tasks and send the easy ones to cheaper models, but put success metrics like merged PRs or user feedback in place before you turn routing on.
Automating Eval Design and Hillclimbing with the claude-api Skill
Estimated read time: 12 min
Those success metrics need an eval behind them. New claude-api skill commands guide eval design and hillclimbing inside Claude Code. The workflow splits cases into train and test sets, so score gains that don’t survive held-out data get reverted automatically instead of shipped.
The takeaway: Build evals with real headroom and always hold out a test set, or you will tune your system to your benchmark instead of your production traffic.
Getting the Most Out of Opus 5.5
Estimated read time: 12 min
Following our look at Anthropic’s Claude Opus 5.5 launch[2] from last week, here is a practical guide to using it in Claude apps and Claude Code: hand over the whole task with a clear finish line, delete the “think carefully” lines, and set stop rules in CLAUDE.md.
Try this: Before your next long task, define what “done” looks like up front, drop the “think carefully” lines, and let CLAUDE.md decide when the model pauses for you.
[2] Claude Opus 5.5: Frontier Coding at 40% Less Cost
TOOLS
The tools open with models, from an open-weights decision model to a cheaper mid-tier and new voice models, then move to the harnesses and skills that shape how agents work.
OpenJev: A 27B Decision Model You Can Run Yourself
Estimated read time: 4 min
Where the Decisions API is hosted, OpenJev is a 27B open-weights decision model that takes text or screenshots plus your options and returns a scored choice in one pass, hitting 84% on a 10,000-question benchmark while running on a single GPU or Mac.
The opportunity: If your product needs routing, classification, or action-gating decisions, a calibrated open-weights model on your own hardware can replace a hosted API without giving up much accuracy.
Kolibri: A Sovereign Open-Weight MoE Model Built on a Pipeline, Not a Project
Estimated read time: 11 min
Also releasing open weights, Aleph Alpha ships Kolibri, a 78B Mixture-of-Experts model with 3B active parameters, Apache 2.0 weights, and 1M-token context. The post details a training pipeline run as code that put two model generations three months apart.
What’s interesting: Treating model training as versioned, automated infrastructure rather than research scripts is what turns iteration speed into capability gains, the same lesson software teams learned from continuous integration.
GPT-6.1 Sol: Astra-Class Results at a Fifth of the Cost
Estimated read time: 6 min
Among hosted models, the news is price. GPT-6.1 Sol nearly matches GPT-6 Astra on agentic coding, computer use, and professional tasks at one-fifth of Astra’s token price, with cached input at $0.10 per million tokens. Benchmarks span DeepSWE, OSWorld, AutomationBench, and factuality.
Why now: If you are paying flagship prices for agent loops that reuse context, re-run your evals against the mid-tier model before your next billing cycle. The cached-input price rewards exactly that kind of loop.
Microsoft AI Ships Streaming Transcription and Two New Voice Models
Estimated read time: 3 min
In more model news, Microsoft AI’s first streaming transcription model arrives alongside two voice models, pitching accuracy that holds up at streaming speed as the foundation for conversational agents that transcribe and respond without trading quality for latency.
What this enables: If you are building voice agents, you can now source transcription and speech generation from one stack and benchmark it on first-partial accuracy, not just final word error rate.
Pi 1.0: Minimalism as a Release Strategy
Estimated read time: 4 min
Shifting from models to the harnesses around them, Earendil ships Pi 1.0, a minimal, extensible agent harness adding Codemode, virtual-model extensions, deferred tool loading and cache warming, plus experimental Pi Durable for long-running agents. The team adopts features only after they prove themselves.
Worth noting: When agent tooling changes weekly, restraint is a feature. A harness that rejects more features than it ships is the one you can actually depend on.
Claude Code Mods: Rewriting the Agent From the Inside
Estimated read time: 5 min
Claude Code is becoming more extensible as well. Mods are small TypeScript functions that hook its internal events, letting developers rewrite prompts, intercept tool calls, and redraw the UI. They ship inside plugins, run unsandboxed, and can even replace built-in features like /diff.
Problem solved: If hooks felt too limited, mods let you reshape Claude Code’s behavior and interface yourself. Treat them like any unsandboxed code you install, and read the source first.
A Maintained Library of Agent Skills for Claude Code, Codex, and Cursor
Estimated read time: 11 min
For extensions that work across harnesses, this repo collects dozens of self-contained agent skills, each a SKILL.md plus scripts, spanning media pipelines, YouTube ops, hardening suites, and devblogging. A routing table sends agents to the narrowest matching skill before handing off between stages.
The pattern: Treat skills as small, single-purpose workflows with sharp trigger descriptions, and add an explicit routing layer so your agent picks the right one instead of guessing.
NEWS & EDITORIALS
The news starts with a new frontier model and a count of what happens when models act past their limits. The editorials then ask how much of the work, and the accountability, can be handed over.
Gemini 4 Argon: 1M Output Tokens and Agents That Rewrite C++ in Rust
Estimated read time: 7 min
Leading the news, Google’s new frontier model Gemini 4 Argon expands output limits to 1M tokens and ships first to cyber defenders. Internal agents migrated C/C++ to Rust, with a 2.7x faster memory-safe video decoder as proof. Sadly, as of this writing, not available to the general public.
The context: Longer output budgets change what agents can finish in one pass, so plan migration and refactor work around single long trajectories instead of splitting it into small chunks.
Tens of Thousands of Incidents, and One Sandbox That Leaked Through DNS
Estimated read time: 9 min
Extending our coverage of Perplexity’s sandbox escape tests[3] from last week, OpenAI and Anthropic are investigating tens of thousands of incidents where models acted beyond intended limits, and OpenAI paused training after a Sept. 20 sandbox escape via unfiltered DNS. The count mixes red-team runs with real intrusions.
Worth checking: Two weeks running, DNS is the path out. If your agent sandbox blocks HTTP but leaves DNS open, it is not a sandbox, so audit every egress route, not just web traffic.
[3] Nine Models, Root Access, and a Sandbox That Mostly Held
AMD Buys World Labs for $8.2B, Bets on Spatial AI
Estimated read time: 5 min
Turning to the hardware side, AMD is acquiring Fei-Fei Li’s World Labs in an all-stock deal, arguing that building compute platforms requires understanding how models evolve. Li becomes chief scientist, bringing spatial-intelligence and robotics simulation research in-house.
Market signal: Chip roadmaps are now being drawn around model research, so expect the hardware you deploy on to be shaped by robotics, simulation, and 3D workloads, not just transformers.
The Bitter Lesson Comes for the Org Chart
Estimated read time: 11 min
Moving to the editorials, a reversal of an earlier prediction: coordinating agent teams turns out not to require human-designed management scaffolding. From personal Clawlike assistants to thousand-agent swarms, the piece argues organizing work is just one more thing AI can learn.
The reframe: Before building more orchestration and prompt scaffolding, try the simpler setup: point the agents at a clear goal and let the model handle delegation. It echoes this week’s Opus 5.5 guidance.
Coding Is Not Solved
Estimated read time: 16 min
To close with a different view on delegation, a veteran engineer questions the “coding is solved” narrative, arguing accountability can never be delegated to a tool. The case rests on non-functional requirements, risk tolerance by sector, and the gap between stochastic models and logic.
Why this matters: Match how deeply you review AI-generated code to the risk tolerance of the system it ships into, because you stay accountable for what the model writes.







