The AI industry keeps betting on bigger models and more GPUs. He made a different bet: make the models we already have run up to 10x cheaper by keeping them from redoing the same work every time.
Our next speaker is Junchen Jiang, co-founder and CEO of Tensormesh and co-creator of LMCache Lab.
His path runs straight through the frontier of AI systems: the Yao Class at Tsinghua, a PhD at Carnegie Mellon, and a computer science professorship at the University of Chicago, where his research zeroed in on a costly inefficiency in how large models handle memory.
That work became LMCache: open-source infrastructure that stores and reuses a model's internal memory (its KV cache) instead of recomputing it from scratch on every request. It cuts inference costs by as much as 10x, and it has spread fast. It now runs inside vLLM and NVIDIA Dynamo, and is used by teams at Bloomberg, Red Hat, Redis, and Tencent.
In 2025 he turned that research into Tensormesh, backed by Valley Capital Partners , NVIDIA, AMD, CoreWeave and Laude Ventures to bring the same optimizations to any team running AI at scale.
Junchen joins the summit through Valley Capital Partners, platinum partner of the AI Summit @ Stanford and a lead backer of Tensormesh. VCP's founding managing partner, Steve O'Hara, who sits on the company's board, calls KV caching "one of the most consequential and underexplored opportunities in AI infrastructure today."
He describes today's AI as a brilliant analyst who reads everything, then forgets it all after each question. He's building the layer that lets it remember.
At the AI Summit @ Stanford, that's exactly the layer he'll dig into. He keynotes "The Next Big Data Layer for AI" and joins leaders from NVIDIA on the panel "The Infrastructure of Intelligence: Building the AI Stack."
This July, he brings that work to Stanford's campus.
📅 July 30 – August 1, 2026
📍 Stanford Campus, Palo Alto
🎟 Secure your place → https://lnkd.in/gJa4DjG8