Blog

Articles

All (93) AI (64) Python (24) LLM (21) Business (17) Infrastructure (16) Machine Learning (14) Investment (12) OpenAI (11)
Part 4. TANREN — an Evolved Scheduler Patched into vLLM Halves Tail Latency Under Heavy Load Series
TANRENEvolutionary SearchvLLMLLMSchedulingPythonGB10

Part 4. TANREN — an Evolved Scheduler Patched into vLLM Halves Tail Latency Under Heavy Load

The scheduler — the heart of LLM serving — evolved by TANREN and A/B tested inside a real vLLM. Across 16 real-trace configurations it beats the default FCFS 15 times (p=2.6e-4), cuts mean/p99 TTFT by 35-49% under heavy load, and does no harm when idle. Includes the honest record of going 0/6 on unseen data and six rounds of diagnosis before the win.

Read more →
Part 3. TANREN — Patching the Evolved Cache Policy into a Real vLLM and Measuring It Series
TANRENEvolutionary SearchvLLMLLMCachingPythonGB10

Part 3. TANREN — Patching the Evolved Cache Policy into a Real vLLM and Measuring It

The cache policy that won in simulation is patched into a live vLLM and measured on real hardware. The simulated win vanished at first; after fixing where and what to measure, tail latency (p99 TTFT) came out 13-16% lower under memory pressure. The prototype's costs and the conditions where it does not help are reported as measured.

Read more →
Part 1. TANREN — Atari Pong 21–0, and the Kill-Shot the Loop Found on Its Own Series
TANRENEvolutionary SearchFunSearchAtariReinforcement LearningLLMPython

Part 1. TANREN — Atari Pong 21–0, and the Kill-Shot the Loop Found on Its Own

How an evolutionary loop — a cheap LLM writing code, scored by a deterministic verifier — reached the theoretical ceiling on Atari Pong (21–0 across 40 games) from nothing but the 128-byte RAM, and discovered a single winning shot without being taught it. Includes the honest protocol and the diagnostic loop that made it work for a few dollars.

Read more →
Is Meta Compute a Sign of an AI Bubble Collapse? Insight
AIAI InfrastructureMetaNeocloudGPU

Is Meta Compute a Sign of an AI Bubble Collapse?

Meta's plan to rent out spare AI compute spooked Neocloud stocks. But looking at pricing, capex, and ad AI, I read Meta Compute not as a sign of compute deflation, but as data centers becoming a rental asset with several revenue routes.

Read more →
Part 2. Measuring KOI on a Text-to-SQL agent — picking a cost-effective model without dropping quality Series
KOIText-to-SQLLLMAI Cost OptimizationOpenRouterNishikiAI Agent

Part 2. Measuring KOI on a Text-to-SQL agent — picking a cost-effective model without dropping quality

Using a Text-to-SQL agent — you ask in plain language, the AI writes SQL and queries an internal database — I measured five models by KOI (KPI ÷ cost). With a quality floor in place, I show how to pick the most cost-effective model, with real numbers and the full procedure.

Read more →
Part 1. VOI and KOI: Metrics for the Token Economy — Maximizing AI's Result per Dollar Series
AITokenmaxxingAI Cost OptimizationModel RoutingKOI

Part 1. VOI and KOI: Metrics for the Token Economy — Maximizing AI's Result per Dollar

Tokenmaxxing is turning AI usage into an internal score, and infrastructure cost is running away. I look at model routers as the first line of defense, their blind spot, and the shift from input-driven to outcome-driven routing with two metrics — VOI and KOI — and an open-source tool, Nishiki.

Read more →
From Semiconductors to Gold, Real Estate, and Subprime: The Financialization and Risk of AI Infrastructure Insight
AIAI InfrastructureFinanceGPULeveraged Loans

From Semiconductors to Gold, Real Estate, and Subprime: The Financialization and Risk of AI Infrastructure

GPUs and ASICs are turning from parts into a store of value, then real estate, then leveraged financial products. I trace CoreWeave, SPVs, and the move into leveraged loans and CLOs, and check the structural similarities to the 2008 subprime crisis with a calm eye on the risk.

Read more →
Reading the AI Era Through Jevons' Paradox: Why Demand Rises and What Talent Will Be Needed Insight
AIJevons ParadoxFuture of WorkSemiconductorsEnergy

Reading the AI Era Through Jevons' Paradox: Why Demand Rises and What Talent Will Be Needed

Does efficiency from AI really destroy demand and jobs? Using Jevons' Paradox from economics, this piece looks at the reversal of demand happening in memory, inference cost, energy, and labor — based on facts.

Read more →
SpaceX IPO and the Question of Valuation: How to Price Starlink's Communications Infrastructure and SpaceXAI Insight
SpaceInfrastructureInvestmentSpaceXStarlinkAI

SpaceX IPO and the Question of Valuation: How to Price Starlink's Communications Infrastructure and SpaceXAI

The reported $1.75 trillion SpaceX IPO is not a rocket-company listing. The real problem is how to price Starlink's cash-generating infrastructure and SpaceXAI — the absorbed xAI business — on a single balance sheet.

Read more →