Topic

AI and agents

Engineering, operating and evaluating AI systems, coding agents and model infrastructure.

Blog posts
137
Focused clusters
02

Focused clusters

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Wavect editorial graphic separating model performance from product defensibility in the Laya vs Jev comparison AI & Agents

Laya vs Jev: What the Benchmarks Mean for AI Startups

Can an open specialist undermine a two-year head start? What Laya actually shows, what the Jev comparison misses, and where a defensible AI business can create value.

Wavect editorial header for Google AX, with a shield representing agent execution boundaries AI & Agents

Google AX Agent Executor: Budgets and Self-Hosting

A sandbox is not a spending cap. What Google AX really controls, what survives suspend/resume, and which dependencies matter when you want to own your AI.

Raw agent traces becoming a persistent wiki and a validated SKILL.md file AI & Agents

WikiSkill: How Agents Evolve SKILL.md from Experience

WikiSkill compiles agent traces into a persistent evidence wiki, then validates atomic changes to SKILL.md. We review the benchmarks, transfer results, limits and production architecture.

Octop connecting users, AI agents, tools and self-hosted data in one application stack AI & Agents

Tencent Octop Review: What Builders Actually Get for Free

Tencent released Octop under MIT, with multi-user agents, a dashboard, automation and storage. We examine what is really included, what is not and what a production fork still costs.

Wigolo fans one AI agent query across search engines, reranks evidence locally and exposes ten web tools through MCP AI & Agents

Wigolo Review: Local Web Intelligence for AI Agents

Wigolo gives AI agents local-first search, fetch, crawl and research without a paid search API. We verify the ten-tool surface, cache, privacy, AGPL license and production limits.

Wavect editorial illustration of a typed decision split for the Jev AI model AI & Agents

Jev AI Review: Decision Models for Agent Workflows

Jev makes typed decisions instead of prose. See how it can route models, govern agent control steps and support bounded automation without replacing hard rules.

Complete article directory

  1. Laya vs Jev: What the Benchmarks Mean for AI Startups
  2. Google AX Agent Executor: Budgets and Self-Hosting
  3. WikiSkill: How Agents Evolve SKILL.md from Experience
  4. Tencent Octop Review: What Builders Actually Get for Free
  5. Wigolo Review: Local Web Intelligence for AI Agents
  6. Jev AI Review: Decision Models for Agent Workflows
  7. Supermemory: AI Agent Memory, RAG and Local Setup
  8. AI Agent Design Patterns: Start Simple, Verify Actions
  9. Claude Mods: Setup, Function Hooks and Security
  10. Voicebox: Local Voice Cloning, Dictation and MCP Setup
  11. Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
  12. OpenAI Agents API Review: Migration, Costs and Data Controls
  13. Spotify shunt Review: Setup, Savings and Limits
  14. OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
  15. SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
  16. Ramp Inspect Architecture 2026: Background Coding Agents at Scale
  17. Model Hardware Standard: Enterprise Guide to Physical AI
  18. Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
  19. Mosaic (YC S26) Review: Shared Memory for Team AI Agents
  20. Ripwire Review 2026: AI Repo Context Without Embeddings?
  21. Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
  22. Phonely Alma Review: Is the Voice LLM Ready for Production?
  23. Utopia Review: Temporal Knowledge Graph for Enterprise
  24. AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
  25. claude-rotate: One Proxy for Multiple Claude Max Accounts
  26. NVIDIA PAIR Review: Local AI Routing and the AMD Gap
  27. Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
  28. Obscura Browser Review: Claims, Limits and Production Fit
  29. Claude Code Design System: 4 Parts for On-Brand UI
  30. Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
  31. PageLM Review: Self-Hosting and Commercial Use
  32. ChatGPT Can Now Log In Without Seeing Your Password
  33. M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
  34. AI Agent Harness, Explained: The Reliability Layer Around an LLM
  35. FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
  36. Darkbloom AI Review: Private Inference on Idle Macs
  37. Ox Alpha Revealed as GLM-5.3-Flash: What Changed
  38. Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
  39. LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
  40. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  41. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  42. OpenViking Review: Filesystem Memory for AI Agents
  43. LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
  44. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  45. TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
  46. Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
  47. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  48. Localized URL Slugs vs hreflang: Why We Keep One English Slug
  49. Can an AI Agent Use Your Product, or Only Read About It?
  50. Graft Review 2026: Do Agent Repo Maps Belong in Git?
  51. How Coding Agents Keep Token Bills in Check with Output Compression
  52. Smarter Token Usage with Your AI Coding Agent
  53. DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
  54. Netflix's vLLM and Triton Stack: 7 Production Lessons
  55. OpenSandbox Review: Is Self-Hosting Worth It?
  56. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  57. Cloudflare Kitesurf Review: Cost, Limits and Production Fit
  58. GitHub Spec Kit Review: Is It Worth the Process?
  59. Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
  60. How to Self-Host LiteLLM in Production: 2026 Guide
  61. AI-Ready Company Wiki: Architecture and Build Guide
  62. Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
  63. Does Claude Watermark Text? The 2026 API Answer
  64. OpenKB Review: Knowledge Compiler vs RAG
  65. Unsloth Desktop Review: A Private Local AI Workstation?
  66. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  67. Firecrawl AnyDoc Review: 14 Formats to Markdown
  68. MCP Cloud vs Manufact Cloud: MCP Hosting Guide
  69. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  70. How to Make AI Writing Sound Human with Agent Skills
  71. NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
  72. OmniRoute AI Routing: Setup and Production Checklist
  73. Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
  74. Hyperagent Review: Cloud AI Agents Without a Server
  75. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  76. Hark Handoff Review: The Agent That Actually Clicks
  77. Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
  78. pdf-inspector Review: Route PDFs Before OCR
  79. PII Redaction Before LLM Prompts: A Practical Pipeline
  80. Agent Reach Review: Costs, Security and Real Limits
  81. Cloudflare Wallets for AI Agents: What Is Live?
  82. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  83. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  84. YC QM Agent Review: Is Quartermaster Ready for Work?
  85. jcode vs Claude Code: Is the Rust Harness Worth Switching To?
  86. Lightpanda Browser for AI Agents: Production Guide
  87. Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
  88. Multi-Model AI Coding Agent Stack: A Team Buying Guide
  89. Is MCP Stateless Now? Your Server Migration Checklist
  90. MCP Is Not a Security Boundary: Protect Agent Data
  91. Fine-Tune Gemma 4 Free with Unsloth and Colab
  92. Can AI Agents Talk to Each Other? A Band Setup Guide
  93. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  94. LeanCTX Technical Field Report: 64.1% Less Context
  95. llmfit Guide: Which Local LLM Fits Your Hardware?
  96. CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
  97. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  98. Enterprise MCP Authorization Architecture
  99. Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
  100. LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
  101. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  102. OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
  103. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  104. Meterless Review 2026: Is This AI Agent Context Layer Ready?
  105. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  106. Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
  107. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  108. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  109. Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
  110. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  111. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  112. AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
  113. T3MP3ST Review 2026: Can It Replace a Penetration Test?
  114. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  115. Open Knowledge Format (OKF): The Enterprise Guide
  116. AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
  117. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  118. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  119. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  120. Cheaper Per Token Can Still Cost More Per Task
  121. How to Use Claude Fable 5.1 in Claude Code
  122. Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
  123. What an Internal AI Assistant Actually Costs in the DACH Region (2026)
  124. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  125. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  126. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  127. The Bottleneck Was Never Intelligence. It Was Context.
  128. ChatGPT Enterprise vs Copilot vs Custom RAG
  129. MCP vs RAG vs Agent Skills vs Custom GPTs
  130. How to Cut LLM Token Costs in 2026
  131. AI Agent Pilot in 30/60/90 Days
  132. Permissions-First RAG over SharePoint, Confluence, Drive
  133. RAG Production-Readiness Checklist for EU Companies
  134. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  135. Why AI Agent Projects Get Cancelled
  136. RAG vs Fine-Tuning vs Long-Context 2026
  137. LLM API Costs 2026. Architecture Shift