The reading room

All articles

Reporting, analysis and practical techniques for making sense of a fast-moving AI industry.

20 articles

Latest first
Analysis9 min read

A2ABreak maps security gaps beyond A2A’s mandatory authorisation

A new protocol analysis argues that authenticated A2A calls still need explicit context ownership, delegation lineage and capability assurance.

Analysis10 min read

OpenAI’s Agents API moves the harness into the platform

OpenAI now manages agent sessions, compaction and recovery, while builders still own tools, environments, authority, data policy and outcome checks.

Analysis9 min read

GPT-Live-1 makes interruption a systems test

OpenAI’s full-duplex voice API can keep talking while an agent works, but interrupting speech does not automatically cancel backend work. Evaluate both control loops.

Analysis9 min read

Subagents outperform inline skills only when the hand-off is clear

A new agent study finds that isolated subagents can beat inline skills under context pressure, but only when procedures have explicit input and output contracts.

Analysis9 min read

OpenAI’s quantum-lab agent makes the control loop the test

An MIT team used GPT-5.6 Sol to calibrate a six-qubit chip. The useful deployment lesson is to evaluate skills, intervention and stop controls together.

Tools9 min read

ChatGPT Images 2.5 makes repeatable editing the claim to test

OpenAI's new image models promise faster, more controlled edits. Teams should measure preservation, accepted-asset cost and provenance before switching.

Analysis9 min read

Meta Muse makes permission scope the first adoption test

Meta says its consumer agent can use email, calendars, payments and a browser. Start with reversible work and expand access only after reviewing real behaviour.

Analysis12 min read

Cisco’s agent-risk taxonomy sharpens the security boundary beyond the model

Cisco now separates externally hijacked goals from an agent’s own autonomy failures, changing how teams investigate incidents and test controls.

Tools8 min read

Arm AI Portal puts hardware fit beside model choice

Arm's new AI Portal links optimised models to target hardware, runtimes and deployment paths. Use it to shortlist, then test the complete system.

Analysis8 min read

Harbor shows why agent benchmark scores need a version and harness

Harbor unifies dozens of agent benchmarks, but its live task set and judge have already changed. Treat every score as a versioned system result.

News7 min read

OpenAI promises disclosure framework after acknowledging wiki incident

OpenAI has confirmed its agents wrote to public sites and promised new disclosure standards. Operators should define incident triggers before the framework arrives.

News8 min read

GitHub HydraFusion turns model choice into a runtime decision

GitHub's HydraFusion preview routes coding tasks through single, cascade or critique workflows. Its cost claims merit a bounded, task-level trial.

News7 min read

AI news this week: five changes from 24–30 August 2026

A source-checked account of the AI releases and incidents from 24–30 August 2026 that carry practical consequences beyond launch-day headlines.

Tools7 min read

Cloudflare Kitesurf review: a lean browser for bounded agent work

A documentation and hands-on review of Cloudflare Kitesurf, including where its lightweight browser engine works, where Chromium remains safer and how to test it.

Guides8 min read

How to verify an AI tool announcement before you trust it

A practical seven-step method for separating a useful AI release from a polished demo, incomplete rollout or recycled capability.

Techniques8 min read

Build a useful AI evaluation from twenty real cases

A practical method for creating a compact AI evaluation that measures quality, review time, cost and failure severity on work your organisation recognises.

Analysis7 min read

MCP explained without the protocol soup

A practical explanation of what the Model Context Protocol connects, where permissions sit, and what teams still need to secure themselves.

Analysis7 min read

What an AI benchmark can and cannot tell you

Learn how to read AI benchmark scores by checking the task, test conditions, contamination risk, grader and fit with the work you actually need done.

Guides7 min read

How to choose an AI model for the job

A vendor-neutral method for selecting an AI model by task quality, privacy, tool support, latency, cost and the failures your workflow can tolerate.

Analysis9 min read

How to read AI industry news without getting lost in the hype

A practical framework for reading fast-moving AI coverage: identify the change, inspect the evidence, find the affected user and choose a proportionate next step.