On CNBC earlier this month, Alex Karp told enterprise buyers to stop handing their data to frontier labs. When you feed proprietary data into someone else's model, you're handing your advantage to a vendor who can turn around and build a competing product. His point: technical buyers should own their compute, their models, and their data stack.
That decision is hiding inside every CX platform. Every AI feature runs on a model. If that model is a frontier API, your data is leaving, your bill scales with volume, and the path from input to output is a black box. Access, pricing, and rate limits can change on someone else's schedule, and when that model sits behind your virtual agent or your QA process, losing it means gaps, brand risk, and churn.
Everyone wants agentic AI. But an agent that acts on its own is only as good as your ability to see what it took in and measure what it did with it. Without access to the inputs, you can't effectively calibrate, and you're shipping a system you can't correct.
That's what Latitude is for. 7 proprietary models fine-tuned for the jobs of CX: transcription, redaction, intent, summarization, inferred CSAT, QA, and Voice of Customer. They run on our own GPU clusters in a dedicated VPC, so customer data never reaches an outside model provider and never trains anyone else's. We own the models; you get full visibility into the inputs and outputs, and we recalibrate against human QA leads every month, because policies, products, and regulations move.
The reason a family of smaller models beats one general-purpose LLM: each job has its own accuracy bar, latency budget, and failure mode. Transcription is judged word by word in noisy audio. Redaction is judged by what it catches before data leaves the boundary. QA is judged by whether a supervisor can trace a score back to evidence. One model can't be tuned for all of that at once. Ours are trained on more than a billion real customer interactions, so they know how support sounds, and each one is pointed at a single task. That's how they hold accuracy on par with GPT, Gemini, and Claude on live CX work, at a fraction of the size.
The result: up to 49x lower cost to serve, up to 3.5x higher throughput, 4x lower latency, running at a 4 trillion token rate with 97% on our own models.
Whose models are your customer conversations training right now?
Level AI
https://lnkd.in/duSWidBR