ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Baseten test: Inference engine written by Claude is 90% faster than vLLM

Baseten test: Inference engine written by Claude is 90% faster than vLLM

AI information • Admin • • 20 views

Can an inference engine be written from scratch by an AI agent and still beat a mature general-purpose framework? In a post updated on October 2, 2026, the Baseten engineering blog reported exactly such a test: drawing on the MetaInfer paper, its engineers had Claude Code (Fable 5) build a custom inference engine from scratch for Qwen-3.6-35B-A3B, running at NVFP4 precision on a single NVIDIA B200, and named it VibeQwen. MetaInfer uses a skills-only framework and requires no model post-training.

The headline numbers: 90% faster single-stream, time to first token cut to less than half

VibeQwen reached 1792 TPS in single-stream decoding, against 943 TPS for a tuned vLLM 0.25.1 — 90% faster. Time to first token was 12ms versus 28ms, a 2.33x improvement, and at concurrency 32, output throughput hit 10307 TPS against vLLM's 6030 TPS, 71% higher. Notably, the baseline was a tuned vLLM, not a default configuration.

The goal was to beat vLLM by 20% across the board without losing accuracy, with BF16 as the accuracy reference. Claude worked largely autonomously for about a week, consuming roughly 1.7 billion tokens (mostly cached input) and about 200 B200 hours, and had already caught up with vLLM within the first few days; when accuracy showed slight numerical differences, human approval was required to proceed.

The second experiment, Sammie: reusing the knowledge base made it faster and cheaper

Baseten then ran a second experiment, Sammie: an agent built an inference service for the image segmentation model SAM 3.1, reaching 91 images per second on a single H100 — 50% higher throughput than Meta's official reference server — in just a few days and about 200 million tokens. It was faster and cheaper because it reused the knowledge base accumulated during the first experiment. The two efforts cost from a few hundred dollars for Sammie to a few thousand dollars for VibeQwen. This reusability of experience may matter more than any single result.

The limits: still experimental, and the accuracy constraint must be locked in first

Both engines are still experimental and carry no production traffic. Baseten also notes that the agent was allowed to reference existing kernels from vLLM and TensorRT-LLM, rather than being fully isolated from source code as in the original paper. This kind of automated optimization rests on two preconditions: inference metrics are quantifiable and correctness can be numerically verified against a full-precision reference implementation, giving every improvement an objective yardstick; and the accuracy constraint must be locked in beforehand — without it, an agent will game the system, for example by outputting the same character over and over to inflate the metrics.

The weakness of general-purpose engines is exactly where custom optimization fits

vLLM, SGLang and TensorRT-LLM win on generality and ease of use, but for a specific combination of model, hardware and workload they are usually not optimal. Inference spending is enormous, and a single-digit percentage improvement is worth millions of dollars at scale — which is precisely where a dedicated engine earns its place.

Recommended Tools

More