vLLM

4 articles tagged with "vLLM"

Part 1. What Makes the New AI Model "Jev" Special? Reproduced with NerveReflex and an Open Model Series

Part 1. What Makes the New AI Model "Jev" Special? Reproduced with NerveReflex and an Open Model

Jev is a new decision-focused AI model announced by ChatGPT co-inventor Diogo Almeida. This post separates the announcement from social media claims, explains why the approach is fast and cheap, and shows how I reproduced the same kind of setup with NerveReflex ideas and the open model Gemma 4 26B.

Read →
Part 2. TANREN — 11 Wins to 1 Against Twenty Years of Cache Classics, on the Future of a Production Trace Series

Part 2. TANREN — 11 Wins to 1 Against Twenty Years of Cache Classics, on the Future of a Production Trace

Cache eviction — deciding what to drop when a cache is full — has been studied for more than twenty years. On that ground, a 53-line evolved Python function beats ARC, LIRS, and W-TinyLFU on future segments of Twitter's production traces. This post covers how the first attempt lost by overfitting to one time period, the two-window training that fixed it, and the public code anyone can rerun.

Read →
Part 4. TANREN — an Evolved Scheduler Patched into vLLM Halves Tail Latency Under Heavy Load Series

Part 4. TANREN — an Evolved Scheduler Patched into vLLM Halves Tail Latency Under Heavy Load

The scheduler — the heart of LLM serving — evolved by TANREN and A/B tested inside a real vLLM. Across 16 real-trace configurations it beats the default FCFS 15 times (p=2.6e-4), cuts mean/p99 TTFT by 35-49% under heavy load, and does no harm when idle. Includes the honest record of going 0/6 on unseen data and six rounds of diagnosis before the win.

Read →
Part 3. TANREN — Patching the Evolved Cache Policy into a Real vLLM and Measuring It Series

Part 3. TANREN — Patching the Evolved Cache Policy into a Real vLLM and Measuring It

The cache policy that won in simulation is patched into a live vLLM and measured on real hardware. The simulated win vanished at first; after fixing where and what to measure, tail latency (p99 TTFT) came out 13-16% lower under memory pressure. The prototype's costs and the conditions where it does not help are reported as measured.

Read →