Cluster

Models and infrastructure

Model selection, inference economics, local deployment, compression and serving architecture.

Blog posts
57
Topic
AI and agents

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Wavect editorial graphic separating model performance from product defensibility in the Laya vs Jev comparison AI & Agents

Laya vs Jev: What the Benchmarks Mean for AI Startups

Can an open specialist undermine a two-year head start? What Laya actually shows, what the Jev comparison misses, and where a defensible AI business can create value.

Wavect editorial illustration of a typed decision split for the Jev AI model AI & Agents

Jev AI Review: Decision Models for Agent Workflows

Jev makes typed decisions instead of prose. See how it can route models, govern agent control steps and support bounded automation without replacing hard rules.

A phone and laptop split one large language model across browser tabs while WebRTC carries activations between them AI & Agents

SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops

A buyer-focused review of SwarmLLM: browser-native 27B inference across phones and laptops, WebGPU/WebRTC architecture, benchmark limits, peer trust, and when to pilot it.

A phone waveform passes through the Alma voice LLM into a production evaluation scorecard AI & Agents

Phonely Alma Review: Is the Voice LLM Ready for Production?

Alma claims 182 ms TTFT, a 206 ms P99 and $0.55 per blended million tokens. See what the benchmark proves, what it omits and how to run a buyer-side voice-agent pilot.

Layered Utopia knowledge records represent preserved history AI & Agents

Utopia Review: Temporal Knowledge Graph for Enterprise

Evaluate Utopia for enterprise knowledge: two-clock history, source provenance, self-hosting limits and a practical pilot plan before you invest.

NVIDIA PAIR routes local AI agent requests across RTX, Mac and an AMD Strix Halo node AI & Agents

NVIDIA PAIR Review: Local AI Routing and the AMD Gap

PAIR routes independent agent requests across local computers without pooling GPUs. See how it works, where Strix Halo falls short and what an AMD port must prove.

Complete article directory

  1. Laya vs Jev: What the Benchmarks Mean for AI Startups
  2. Jev AI Review: Decision Models for Agent Workflows
  3. SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
  4. Phonely Alma Review: Is the Voice LLM Ready for Production?
  5. Utopia Review: Temporal Knowledge Graph for Enterprise
  6. NVIDIA PAIR Review: Local AI Routing and the AMD Gap
  7. Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
  8. PageLM Review: Self-Hosting and Commercial Use
  9. M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
  10. FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
  11. Darkbloom AI Review: Private Inference on Idle Macs
  12. Ox Alpha Revealed as GLM-5.3-Flash: What Changed
  13. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  14. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  15. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  16. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  17. Netflix's vLLM and Triton Stack: 7 Production Lessons
  18. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  19. How to Self-Host LiteLLM in Production: 2026 Guide
  20. AI-Ready Company Wiki: Architecture and Build Guide
  21. Does Claude Watermark Text? The 2026 API Answer
  22. OpenKB Review: Knowledge Compiler vs RAG
  23. Unsloth Desktop Review: A Private Local AI Workstation?
  24. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  25. Firecrawl AnyDoc Review: 14 Formats to Markdown
  26. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  27. OmniRoute AI Routing: Setup and Production Checklist
  28. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  29. pdf-inspector Review: Route PDFs Before OCR
  30. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  31. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  32. Fine-Tune Gemma 4 Free with Unsloth and Colab
  33. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  34. llmfit Guide: Which Local LLM Fits Your Hardware?
  35. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  36. Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
  37. LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
  38. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  39. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  40. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  41. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  42. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  43. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  44. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  45. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  46. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  47. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  48. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  49. Cheaper Per Token Can Still Cost More Per Task
  50. Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
  51. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  52. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  53. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  54. How to Cut LLM Token Costs in 2026
  55. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  56. RAG vs Fine-Tuning vs Long-Context 2026
  57. LLM API Costs 2026. Architecture Shift