Fireworks AI’s cover photo
Fireworks AI

Fireworks AI

Software Development

San Mateo, CA 50,085 followers

The frontier platform for training and inference on open-weights models at scale.

About us

Fireworks is the fastest way to build, tune, and scale AI on open models. Ship production-ready AI in seconds on our globally distributed cloud infrastructure, optimized for your use case. Fireworks powers production workloads at companies like Uber, Doordash, Notion, and Cursor—delivering 15× faster speed, 4× lower latency, and 4× more concurrency than closed models.

Website
https://fireworks.ai
Industry
Software Development
Company size
51-200 employees
Headquarters
San Mateo, CA
Type
Privately Held
Founded
2022
Specialties
LLMs, Generative AI, artificial intelligence, developer tools, software engineering, and inference

Locations

Employees at Fireworks AI

Updates

  • Building upon our past work with Trilogy AI Center of Excellence, it's clear that the unit economics of AI matter deeply. Trilogy is not just a day-zero adopter of frontier open models, but built a playbook for cyber defense using Kimi K3 on Fireworks. Audit every repo, triage every scan, review every PR. Get the full details: https://lnkd.in/gf-9winp

    View organization page for Fireworks AI

    50,085 followers

    The new reality for teams is model routing. Fireworks is being used inside early agent workflows where model selection happens dynamically based on cost and performance. Two perspectives from our partners at Trilogy make this concrete: • Adoption → why they moved to Fireworks (cost pressure, rate limits, flexibility) • Usage → how it’s integrated into their systems (provider abstraction + orchestration like Open Symphony)

  • Fireworks AI reposted this

    I'm super excited to share we launched Fireworks Nexus! When I give talks, I listen to the audience, trying to understand and track their reaction in the moment and over time. When I raised "scaling AI into bankruptcy", the room was laughing two years ago. Now the room goes silent. What used to feel like someone else’s problem has become their acute pain. Historically, most teams adopted the best frontier models because that was the right decision. Today, open models have caught up, and most routine development tasks no longer require the most expensive frontier model. But now you have a different problem: which model should run which task? Switching models, routing requests, tuning caching, managing enterprise controls, and figuring out where the AI budget is actually going quickly becomes an infrastructure and operation problem of its own. That’s why we built Fireworks Nexus. Nexus is an intelligence layer that sits underneath the agentic harnesses developers already use. It gives engineering teams access to the latest open models, including GLM 5.2 and Kimi K3, automatically routes requests to the right model, maximizes cache efficiency, enforces zero data retention, and provides a single view into how AI is being used across the organization. We recently published a study on oracle routing showing how much this can matter in practice [https://lnkd.in/erifJ24V] Teams should measure whether an open model actually delivers the same outcome as a frontier model on their own codebase and business workflows, rather than relying on generic benchmarks or intuition. I don’t think the future is every developer doing tedious work in model selection. The future is an intelligence layer underneath every organizations – continuously incorporating the best open models as they emerge while keeping quality high, operational complexity invisible, and AI spend sustainable. Two years ago, the problem was whether AI could create enough value. Today, the problem is making sure that value doesn’t get eaten by the cost of running it.

    View organization page for Fireworks AI

    50,085 followers

    We ran Kimi K3 against Fable on ~1,000 agentic tasks, and we learned that open source models are specialists. Split the results by task type and you see: → K3 outperformed on security, crypto, and long terminal loops. → Fable took multi-language breadth and web/data viz. They complement one another. Predictively routing each task to whichever model handles it best achieved 93% accuracy, above either model by itself, and at much lower cost than running Fable alone. The ideal theoretical router sends 72-96% of traffic to the open model. At scale, this means the closed model becomes the exception you reach for instead of the default. The routing layer is The Thing now. Model quality is important, don't get us wrong. But quality is table stakes; the beginning of the story. Kimi K3 lands with open weights on Fireworks July 27. Courtesy of Kimi (Moonshot AI) See the full breakdown → https://lnkd.in/gi_8gNqA

    • No alternative text description for this image
  • Notion's webinar (tomorrow morning at 10am PT!) will feature our recently launched Fireworks Nexus as a way to help teams build with AI in a way that feels useful, sustainable, and ready to scale. Sarah Sachs, AI Engineering Lead at Notion, will share how to think about AI systems, budgets, and frameworks in the real world, including what teams should consider to avoid going from AI-pilled to AI-poor.

    View organization page for Notion

    1,085,938 followers

    AI experiments are easy to start. Scaling them responsibly is harder. Join Notion’s Sarah Sachs, AI Engineering Lead, for a practical webinar on how teams can make smarter AI build decisions, manage budget tradeoffs, and stay flexible across tools and models.

  • View organization page for Fireworks AI

    50,085 followers

    ICYMI yesterday, Kimi K3 is live and Fast on Fireworks. Things to know: - Performance: same quality as top models from Open AI and Anthropic on coding and agent tasks, at a fraction of the cost. - Privacy: Fully US-hosted endpoints and Zero Data Retention, making it day-zero compliant for large enterprises. - Serverless + Training: no GPUs to manage, simply scale instantly with Fireworks Serverless (Fast, Priority tiers available) or fine-tune K3 on your own data using our Serverless LoRA training supported. Stop renting intelligence. We are excited to see what you build. Get started with Kimi K3 now: https://lnkd.in/gRw_Hg4V

  • Fireworks AI reposted this

    View organization page for Mercor

    799,681 followers

    We got Day-0 access to Kimi K3 via Fireworks AI, and it’s now the highest scoring open weights model on APEX-Agents. Mean score: 55.4% Pass@1: 39.3% Rank: #7 overall K3’s low token usage paired with their lower pricing makes it a compelling combination. It uses just over half the input tokens as Opus 5 while costing 40% less per token. That represents a significant potential cost savings. We first evaluated Kimi K2.5 in February at 14.4% Pass@1. Since then the family has climbed steadily: K2.6 reached 18.9%, K2.7 Code reached 27.6%, and K3 now sits at 39.3%. Pass@1 nearly tripled in six months, and the frontier is in sight: the current APEX-Agents leader, Opus 5 (Max), scores 43.5% Pass@1 and 60.6% mean. APEX-Agents includes 480 long-horizon tasks across professional services domains like investment banking, management consulting, and corporate law. Across the three APEX-Agents domains, K3 scored: Management consulting: 44.4% Pass@1 (#6) Investment banking: 40.5% Pass@1 (tied #5 with GPT-5.5 xHigh) Corporate law: 33.1% Pass@1 (#6) In corporate law, K3's mean score runs 25+ points ahead of its Pass@1. That means K3 does much of a legal task correctly. It can find the right documents and draft usable language, but is not yet ready to hand over a finished, client-ready document. That profile makes it a strong first-pass assistant for legal work, with a lawyer still closing the last mile. Running APEX-Agents on an open weights model at release requires an inference partner like Fireworks AI. Their zero data retention endpoint lets us run evals with no risk of task data leaving the environment, and their serverless API scales to what long-horizon agentic evals demand without any of the headache of managing the underlying serving infrastructure. Congratulations to the Kimi (Moonshot AI) team!

    • No alternative text description for this image
  • Fireworks AI reposted this

    What a week of frontier intelligence with Kimi K3 and Opus 5 launch! We compared these two great models on a few benchmarks of the quality and cost, based on our most common use cases. Important measurement is **task-base cost**, not token-base cost. In general, open models tend to be more verbose than close models. Hence the quick study. Across SWE (480), Algorithmic(100), Terminal (83), the per-task quality is close with Opus 5 being on par or better, and K3's per-task cost is 2x to 4.6x cheaper, based on the serverless pricing. We aim to increase per-task serving efficiency (via Fireworks Inference) and quality (via Fireworks Training). Share what you find out for the tasks you care about. 👇 More details of our study -- https://lnkd.in/eYtHQR2D

    • No alternative text description for this image
  • Today, we’re introducing Fireworks Nexus. Fireworks Nexus gives engineering organizations control over the intelligence powering the AI tools they already use. Instead of being locked into a single provider’s pricing, models, and roadmap, you can choose the best model for every task, manage spend centrally, and continuously measure quality as new models emerge. Now you can get the same or better engineering outcomes while spending less. Learn more at the link in the comments.

Similar pages

Browse jobs