Agent engineering
Coding agents, MCP, context systems, evaluation and the controls required for dependable automation.
- Blog posts
- 80
- Topic
- AI and agents
Start with the cornerstone
Get the next AI and agents field note
One concise email when we publish. No tracking pixels, and no inbox filler.
Latest in this collection
Google AX Agent Executor: Budgets and Self-Hosting
A sandbox is not a spending cap. What Google AX really controls, what survives suspend/resume, and which dependencies matter when you want to own your AI.
WikiSkill: How Agents Evolve SKILL.md from Experience
WikiSkill compiles agent traces into a persistent evidence wiki, then validates atomic changes to SKILL.md. We review the benchmarks, transfer results, limits and production architecture.
Tencent Octop Review: What Builders Actually Get for Free
Tencent released Octop under MIT, with multi-user agents, a dashboard, automation and storage. We examine what is really included, what is not and what a production fork still costs.
Wigolo Review: Local Web Intelligence for AI Agents
Wigolo gives AI agents local-first search, fetch, crawl and research without a paid search API. We verify the ten-tool surface, cache, privacy, AGPL license and production limits.
Supermemory: AI Agent Memory, RAG and Local Setup
Give your AI agent context between sessions without building the memory pipeline yourself. A practical Supermemory guide with API code, local setup and an honest benchmark review.
AI Agent Design Patterns: Start Simple, Verify Actions
More loops are not automatically better. Choose the smallest agent architecture that solves the failure, and separate a proposed action from permission to execute it.
Complete article directory
- Google AX Agent Executor: Budgets and Self-Hosting
- WikiSkill: How Agents Evolve SKILL.md from Experience
- Tencent Octop Review: What Builders Actually Get for Free
- Wigolo Review: Local Web Intelligence for AI Agents
- Supermemory: AI Agent Memory, RAG and Local Setup
- AI Agent Design Patterns: Start Simple, Verify Actions
- Claude Mods: Setup, Function Hooks and Security
- Voicebox: Local Voice Cloning, Dictation and MCP Setup
- Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
- OpenAI Agents API Review: Migration, Costs and Data Controls
- Spotify shunt Review: Setup, Savings and Limits
- OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
- Ramp Inspect Architecture 2026: Background Coding Agents at Scale
- Model Hardware Standard: Enterprise Guide to Physical AI
- Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
- Mosaic (YC S26) Review: Shared Memory for Team AI Agents
- Ripwire Review 2026: AI Repo Context Without Embeddings?
- Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
- AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
- claude-rotate: One Proxy for Multiple Claude Max Accounts
- Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
- Obscura Browser Review: Claims, Limits and Production Fit
- Claude Code Design System: 4 Parts for On-Brand UI
- ChatGPT Can Now Log In Without Seeing Your Password
- AI Agent Harness, Explained: The Reliability Layer Around an LLM
- Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
- LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
- OpenViking Review: Filesystem Memory for AI Agents
- LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
- TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
- Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
- Localized URL Slugs vs hreflang: Why We Keep One English Slug
- Can an AI Agent Use Your Product, or Only Read About It?
- Graft Review 2026: Do Agent Repo Maps Belong in Git?
- How Coding Agents Keep Token Bills in Check with Output Compression
- Smarter Token Usage with Your AI Coding Agent
- DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
- OpenSandbox Review: Is Self-Hosting Worth It?
- Cloudflare Kitesurf Review: Cost, Limits and Production Fit
- GitHub Spec Kit Review: Is It Worth the Process?
- Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
- Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
- MCP Cloud vs Manufact Cloud: MCP Hosting Guide
- How to Make AI Writing Sound Human with Agent Skills
- NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
- Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
- Hyperagent Review: Cloud AI Agents Without a Server
- Hark Handoff Review: The Agent That Actually Clicks
- Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
- PII Redaction Before LLM Prompts: A Practical Pipeline
- Agent Reach Review: Costs, Security and Real Limits
- Cloudflare Wallets for AI Agents: What Is Live?
- YC QM Agent Review: Is Quartermaster Ready for Work?
- jcode vs Claude Code: Is the Rust Harness Worth Switching To?
- Lightpanda Browser for AI Agents: Production Guide
- Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
- Multi-Model AI Coding Agent Stack: A Team Buying Guide
- Is MCP Stateless Now? Your Server Migration Checklist
- MCP Is Not a Security Boundary: Protect Agent Data
- Can AI Agents Talk to Each Other? A Band Setup Guide
- LeanCTX Technical Field Report: 64.1% Less Context
- CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
- Enterprise MCP Authorization Architecture
- OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
- Meterless Review 2026: Is This AI Agent Context Layer Ready?
- Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
- Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
- AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
- T3MP3ST Review 2026: Can It Replace a Penetration Test?
- Open Knowledge Format (OKF): The Enterprise Guide
- AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
- How to Use Claude Fable 5.1 in Claude Code
- What an Internal AI Assistant Actually Costs in the DACH Region (2026)
- The Bottleneck Was Never Intelligence. It Was Context.
- ChatGPT Enterprise vs Copilot vs Custom RAG
- MCP vs RAG vs Agent Skills vs Custom GPTs
- AI Agent Pilot in 30/60/90 Days
- Permissions-First RAG over SharePoint, Confluence, Drive
- RAG Production-Readiness Checklist for EU Companies
- Why AI Agent Projects Get Cancelled