Insights
Field notes, essays, and analysis exploring AI-native systems, logistics intelligence, and modern enterprise architecture.
Are We the Horse or the Cart? Employment After AI's Exponential Curve
The last general-purpose technology this fast — the car — net-created work for humans but retired the horse as labor entirely. Which precedent AI follows is the real question behind Musk's "universal high income" and a 20x disagreement among serious forecasters.
Whose Safety Are You Buying? Security and Governance Across the Big AI Labs
"AI safety" is three questions wearing one name — the frontier-risk framework a lab imposes on itself, the closed-vs-open access architecture it ships, and the data defaults on the tier you actually use. Anthropic, OpenAI, Google, and Meta compared across all three.
The Decision-Only Model: What Jev Signals About How AI Gets Used Next
TypeSafe AI's Jev came out of stealth generating no text at all — it returns typed decisions with calibrated probabilities instead. It points at a split in how production AI is built: a fast decision layer beneath the generative one.
The Accountability Gap: Governance Frameworks Meet the Autonomous Agent
NIST's AI RMF, ISO 42001, and the EU AI Act govern the model. Autonomous agents act as non-human identities with standing access — and the owner of record, the decision-level audit trail, and the kill switch are all still being retrofitted.
When Agents Can't Forget: What LedgerBench Found About Requirement Memory
A pre-registered benchmark on how AI coding agents remember requirements across sessions: an append-only test ledger roughly doubles retention, but calcifies into stale checks that coerce wrong edits — and an isolated judge recovers the benefit at 56% of the cost.

The SLM Default: Why 2026's Production Agents Run Small Models First
Frontier launches still make the headlines, but the agent stacks actually shipping this year default most steps to a small, fine-tuned model and escalate to a frontier one only when a step earns it.

The Allowlist Illusion: Why Command Approval Keeps Failing in Coding Agents
Three unrelated 2026 disclosures — Cursor, Semantic Kernel, and the wider prompt-injection numbers behind them — converge on the same gap: an allowlist checks what a command looks like, not what put it there.

Stale by Default: Why Agents Act on Superseded Data
Retrieval systems rank by similarity, and a revised policy clause sits almost on top of the version it replaced. Temporal validity belongs in the metadata filter, not in the ranker.

MCP in Production: What It Actually Takes to Ship Reliable AI Agents
The Model Context Protocol solved the tool-integration problem. Reliability — scoped access, versioned contracts, idempotent writes, full observability — is still the part teams have to build themselves.

Why RAG Fails in Production — and What to Do About It
Retrieval-augmented generation works remarkably well in demos. Operational environments are a different problem entirely. Real enterprise data is messy by nature.

Fine-tuning vs. Prompting — The Real Tradeoff
The debate between fine-tuning and prompt engineering isn't just technical — it's an operational decision. Here is a guide on where the trade-off actually lies.

Text-to-SQL for Operational Analytics — Beyond the Toy Examples
Making natural language querying work against real freight and procurement data requires hybrid search, metadata filters, self-correction loops, and context budgeting.

LLMOps — What Enterprise Teams Miss When Moving to Production
Deploying a prototype is straightforward. Operating one in production requires observability, prompt versioning, structured evaluation frameworks, and context window discipline.

From Dashboards to Intelligence Systems
Why visualizing data is no longer enough — and what comes after the dashboard era.

Building AI Procurement Intelligence Systems
Procurement workflows are fragmented by design. RFQs arrive as spreadsheets, PDFs, emails, pricing tables, carrier notes, and operational updates — usually spread across disconnected systems.
