FROM THE TRENCHES Issue №244

Opinions, Not Press Releases

Browse a curated front page, then move through focused topic and cluster hubs without losing older work to chronology.

Explore by topic
06
Focused clusters
12
Blog posts
244

Latest articles

Nine articles per page, ordered by publication date.

A voice waveform splitting between a managed Miso TTS API and a private self-hosted GPU AI & Agents

Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality

Miso TTS is expressive and open-weight, but the 110 ms claim belongs to its hosted H100 API, not local inference. Compare hardware, pricing, licensing and production risk before choosing.

An MCP client presents an audience-bound token to a resource server that validates it against an identity provider before any tool runs AI & Agents

Enterprise MCP Authorization Architecture

OAuth 2.1, audience-bound tokens and no token passthrough are the spec's hard rules. The rest, multi-tenant isolation, per-tool scopes, delegated access and audit, is a reference design. Here is the whole trust boundary in two diagrams.

EU AI Act Article 50 transparency obligations mapped to a build checklist for SaaS chatbots, AI agents and generated content Business & Regulation

EU AI Act Article 50 Checklist for SaaS and AI Agents

Article 50 has applied since 2 August 2026. Audit chatbot and agent disclosures, machine-readable marking, deepfake labels, public-interest text, logging, and the evidence kept for enforcement.

Data-oblivious vector quantization turning float32 embeddings into raw 2-bit codes with 16x fewer bytes AI & Agents

Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?

TurboQuant now ships in Qdrant 1.18, while TurboVec has reached 1.0. Learn what the 16x figure applies to, how current benchmarks compare, and what to test.

Eight isolated per-rank KV caches on one server collapsing into a single shared host-side cache layer that every serving process can read AI & Agents

LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache

LMCache reported mean time-to-first-token falling from 3.98s to 0.29s in one multi-turn Qwen3-235B benchmark. Here is what the result shows, what it does not prove, and how to test reuse on your own workload.

Lossless BF16 weight compression compared with 8-bit GGUF quantization across exactness, runtime evidence and production readiness AI & Agents

Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?

A fact-checked decision guide to the GLM-5.2 result, exact BF16 decoding, Q8 quantization, prior systems, runtime evidence, fleet economics and a 14-day pilot.

Paid pilot, proof of concept and design partner compared by the commercial evidence each produces Product & MVP

Paid Pilot vs PoC vs Design Partner: Which One Proves Demand?

Decision guide for choosing the right evidence test: technical feasibility, customer learning or willingness to pay, with a seven-condition paid-pilot scorecard and contract brief.

Nested AI agent evaluation sandbox with isolated compute, controlled egress, ephemeral credentials and out-of-band monitoring AI & Agents

OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent

A fact-checked incident analysis and procurement checklist for nested isolation, egress, credentials, control-plane separation, benchmark integrity, monitoring and forensic fallback.

Parallel AI coding agents use isolated Git worktrees or Jujutsu workspaces backed by one shared repository store Delivery & QA

Git Worktrees vs Jujutsu for AI Coding Agents: 2026 Decision Guide

Buyer guide to repository caching, partial clone, Git worktrees and Jujutsu workspaces, with compatibility gaps, cost metrics and a measurable 14-day pilot.

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.