OpenAI's best new model costs $30 per million output tokens. A Chinese alternative costs 30 cents. “Pricing gap” doesn’t do that justice - it’s more like two different universes. When the price gap was just about helpful chatbots, nobody really cared. A few questions here and there; the cost difference is a rounding error. But the way companies use AI has fundamentally changed since ~2023. Programming and agentic tasks went from about 11% of OpenRouter volume to more than half. An overnight coding run can call a model 1000s of times. Goldman Sachs analysts have projected that the spread of agents will drive a 24-fold increase in token consumption by 2030. At that scale, per-token price stops being a line item and becomes the budget itself. The benchmarking firm Artificial Analysis runs every major model through the same 10 evaluations and totals the cost. Anthropic's Claude came in at $4,811. OpenAI's at $3,357. DeepSeek at $1,071. Moonshot's Kimi at $948. And Zhipu's GLM at $544. The premium American model costs nearly nine times as much as the cheapest Chinese one for the same work. Figma CEO Dylan Field described three phases of corporate AI like this: first, nobody uses it, then everyone must, with teams literally holding competitions of who can spend the most with tokens; then, third, the morning-after realization that everyone is spending too much. Sundar Pichai told his own developer conference that many companies are already blowing through their annual token budgets … and it was only May. When enterprises hit that third phase and agents multiply their token usage by 24x, a 70x price gap becomes the sole deciding factor (to say nothing of the access bans; that’s an entire other topic). P.S. For the framework underneath all of this, see my article How a 30-Year-Old Chart Explains Intel's Fall, Nvidia's Rise, and DeepSeek's Disruption, which lays out why good enough and cheap beats best and gated with predictable regularity. Link in the comments.
AI Model Training Costs and Tariffs
Explore top LinkedIn content from expert professionals.
Summary
AI model training costs and tariffs refer to the expenses associated with developing and running artificial intelligence models, including extra charges based on language, data, or usage patterns. As AI adoption grows, understanding these costs is crucial because prices vary widely depending on the model, language used, and specific needs of an organization.
- Compare pricing structures: Evaluate several AI models to find the best balance between performance, cost, and specific task requirements, as prices can differ by up to seventy times depending on the provider and features.
- Watch for language tariffs: Using AI in non-English languages can increase costs significantly, so consider models optimized for your preferred language or invest in local solutions to reduce these expenses.
- Account for usage patterns: High-volume or complex workflows can rapidly increase costs, so review pricing details carefully and choose models that match your actual usage and budget constraints.
-
-
Tariffs have been a hot topic recently but did we miss a massive “AI Tariff” hidden in plain sight? If you use AI in Europe in a European language, it can cost 2x compared to speaking to AI in English. Why? Because today's AI is not optimized for European languages, but for English 🇺🇸 and Chinese 🇨🇳 . A good example of why its important for European countries to become AI Shapers, instead of AI Takers. I did a quick analysis (done on a flight Europe 🇪🇺 → Brazil 🇧🇷 ) and compared a sample of random sentences in English vs several European languages and measured the actual AI token tariff for some European languages for American, Chinese and European models: 🇳🇱 Netherlands: +124% 🇫🇷 France: +106% 🇪🇸 Spain: +84% 🇵🇱 Poland: +151% 🇵🇹 Portuguese & 🇧🇷 Brasil: +94% 👉 Example question asked to AI in English: What field of scientific research, focusing on DNA, genes, and cellular processes, is providing innovative treatment options and therapeutic strategies for patients suffering from uncommon disorders and complicated health conditions? Answer: Molecular biology → 41 tokens on average Same Question in Dutch 🇳🇱: Welk gebied van wetenschappelijk onderzoek, gericht op DNA, genen en cellulaire processen, biedt innovatieve behandelingsopties en therapeutische strategiën voor patiënten die lijden aan ongewone aandoeningen en gecompliceerde gezondheidsproblemen? Antwoord: Moleculaire biologie → 77 tokens on average That’s 87% more tokens to answer the same question only because it is in Dutch‼️ And since tokens = cost = productivity tax, that’s a 87% more expensive to answer the same question 🤯. And this happens across many languages... As AI becomes the biggest productivity driver of this era, a ‘token tariff’ on local languages seems like a big strategic disadvantage for every non-English economy. The China solution 🇨🇳 The Chinese did something smart: they invested early in native tokenizers & models, trained on Chinese data. Result: they eliminated a ~51% “tariff” when using Western models, and even gained a 3–7% discount when using Chinese in Chinese-built models (e.g. DeepSeek, Qwen and many others). That’s a structural cost advantage for an entire language economy. If you’re a Minister of Finance, Economics in a non-English-speaking country, would it not make sense to actively find ways to stimulate local AI (datasets, companies, etc) to remove a structural productivity tax? The entire economy benefits from the resulting productivity discount. Or am I missing something here? PS: I didn’t analyze every European language (yet), but I’d bet the AI token tariff is even higher for less-spoken ones: 🇬🇷 Greek, 🇸🇪 Swedish, 🇩🇰 Danish, 🇳🇴 Norwegian, 🇧🇬 Bulgarian, etc. PPS: This isn’t just about cost — higher token counts also mean: more latency, larger context windows consumed, higher energy usage per query ⚡
-
$7,225 for one day of coding. And Cursor isn't even the worst example. Replit's margins went negative. Anthropic throttles its best users. I mapped pricing across 50 AI startups. Six distinct patterns emerged. The core tension: traditional SaaS has near-zero marginal cost per user. AI products pay for compute on every interaction. A casual Claude user costs pennies. A developer running Claude Code all day costs tens of thousands per month. Your best users are your most expensive users. That tension is breaking every pricing model in the market. Cursor charged a flat 500 requests/month. Worked fine until users leaned into multi-step agent workflows. They switched to credit pools. One developer burned 500 requests in a single day. The plan description changed from "Unlimited" to "Extended" twelve days after launch. Replit grew 15x in ten months ($16M to $252M ARR). But they were buying revenue with compute. When they launched a more autonomous agent, margins crashed to negative 14%. They had to invent "effort-based pricing" mid-flight. Anthropic played it differently. Their $17/$100/$200 tiers map to genuinely different user personas, not volume bands. A casual user and a Claude Code developer are different products with different willingness to pay. The lesson across all 50 companies: before you set any price, pull the cost distribution. What does your P10 user cost? P50? P90? If the ratio exceeds 10x, flat pricing will break. In AI products, it almost always exceeds 10x. Full guide with all 6 models, 4 case studies, and a decision tree: https://lnkd.in/gdKaQSMk
-
#MIT's new "Radial Attention" makes Generative Video 4.4x cheaper to train and 3.7x faster to run. Here's why: The problem with current AI video? It's BRUTALLY expensive. Every frame must "pay attention" to every other frame. With thousands of frames, costs explode exponentially. Training one model? $100K+ Running it? Painfully slow. Massachusetts Institute of Technology, NVIDIA, Princeton, UC Berkeley, Stanford, and First Intelligence just changed the game. Their breakthrough insight: Video attention works like physics. - Sound gets quieter with distance - Light dims as it travels - Heat dissipates over space Turns out, AI video tokens follow the same rules. Why waste compute power on distant, irrelevant connections? Enter Radial Attention: Instead of checking EVERY connection: • Nearby frames → full attention • Distant frames → sparse attention • Computation scales logarithmically, not quadratically Technical result: O(n log n) vs O(n²) Translation: MASSIVE efficiency gains Real-world results on production models: 📊 HunyuanVideo (Tencent): • 2.78x training speedup • 2.35x inference speedup 📊 Mochi 1: • 1.78x training speedup • 1.63x inference speedup Quality? Maintained or IMPROVED. What this unlocks: 4x longer videos, same resources 4.4x cheaper training costs 3.7x faster generation Works with existing models (no retraining!) And, MIT open-sourced everything: https://lnkd.in/gETYw8eT The bigger picture: The internet is transforming. BEFORE: A place to store videos from the real world NOW: A machine that generates synthetic content on demand Think about it: • TikTok filled with AI-generated content • YouTube creators using AI for entire videos • Streaming services producing personalized shows • Educational content generated for each student This changes everything. Remember when only big tech could afford image AI? 2020: GPT-3 → Only OpenAI 2022: Stable Diffusion → Everyone 2024: Midjourney everywhere Video AI is next. Radial Attention probably just accelerated the timeline. The future isn't coming. It's here. And it's more accessible than ever. Want to ride this wave? → Follow me for weekly AI breakthroughs → Share if this opened your eyes → Try the code: https://lnkd.in/gETYw8eT What will YOU create when video AI costs 4x less? #AI #VideoGeneration #MachineLearning #TechInnovation #FutureOfContent
-
People debate restaurants for 45 mins but pick their AI models in 45 secs. 300+ models. Millions at stake. Yet people pick the one who has the best marketing budget. 67% of organisations globally now run LLM-powered workflows. API pricing spans from $0.15 to $60+ per million tokens, depending on what you pick. 239+ models are being actively evaluated on major benchmarks as of 2026. And the model that topped every leaderboard six months ago is already in the middle of the pack. So here's the decision framework and where today's most relevant models actually sit within it 👇 🔐 STEP 1: Local or Cloud? Sensitive data? Go local. And local no longer means weak. → Low/medium tasks (extraction, simple logic): IBM Granite 4.1 (3B, free), MiniCPM-V 4.6 (1.3B, free + multimodal), NVIDIA Nemotron-4 15B → Complex local reasoning: DeepSeek-V4-Flash (self-hostable, 13B active parameters) paired with Chain-of-Thought prompting does heavy lifting without touching a cloud API. Cloud-connected? Horsepower isn't the problem. Cost and latency are. ⚡ STEP 2: Capability vs. Speed/Cost? Maximum intelligence needed: → Claude Mythos Preview leads on GPQA reasoning. Gemini 3.1 Pro is the strongest for coding agent workflows. GPT-5.5 hits 74.9% on SWE-Bench for software engineering tasks. Speed and cost matter more: → GPT-5.5 Instant is now the default ChatGPT model — fast and affordable. → Qwen3-Max at ~$1.20/M tokens matches Claude Sonnet 4.6 on Arena-Hard. DeepSeek V3.2 sits at $0.28/M input — arguably the best cost-performance ratio available right now. 🎯 STEP 3: Does a specialised model exist for your domain? Almost always yes — and it almost always wins on narrow tasks. → Code agents: GPT-5.3 Codex (tuned for agentic IDE workflows at $1.25/M), DeepSeek-V4-Pro, Gemini 3.1 Pro → Budget-conscious general use: Kimi K2.6 at $0.95/M is the cheapest model in the current top 10 by GPQA Diamond. Gemini Flash for real-time agentic use cases. The shift nobody's pricing into their stack yet: GPT-4-level performance cost $30/M tokens in 2023. Today, you get equivalent capability for under $1/M. Intelligence is commoditising faster than teams are adapting. The open-source ecosystem is now self-hosting models that would have required expensive API access just a year ago. The right model isn't the most powerful one. It's the one that fits your task, your data privacy requirements, and your budget simultaneously. Every time. Saved the full decision tree above as your cheat sheet. Stop choosing models on vibes. Start choosing a strategy. If you liked this, you'll hate how much time you've been wasting on the wrong models. Follow me. I break down AI decisions that actually matter for builders.
-
I spent six months investigating Retrieval-Augmented Generation economics. What I found will make you reconsider every "cheaper AI" pitch you've ever heard. The promise: Build intelligent systems for $15,000 instead of $200,000. Democratize AI. Level the playing field. The reality I documented: → 72% of enterprise RAG projects fail within 12 months → Actual costs hit $1.5 million (10x vendor quotes) → Kenyan workers labeling your training data earn $1.32/hour reviewing 700 traumatic cases per shift → Venezuelan annotators make $0.11-$0.90 hourly → 185 workers unionized in Kenya. All were terminated. While Silicon Valley celebrates Pinecone's $750 million valuation (107x revenue), I followed the money to its source. I found Chen, a healthcare director whose $15K RAG pilot became a $1.2M terminated project. I found Maria in São Paulo, watching her journalism get embedded in corporate systems without a cent in licensing fees. I found workers across Kenya, Venezuela, Philippines, Syria, Bulgaria, Argentina, Ghana, and Colombia. 18-20 hour workdays. PTSD from content moderation. NDAs silencing their testimony. Productivity bonuses worth 50% of wages incentivizing them to process suicide videos in 50 seconds. The data centers powering these systems? They'll consume 1,000 TWh by 2026, equal to Japan's entire electricity usage. Ireland's data centers already take 17% of national power. Nevada facilities compete with Pyramid Lake Paiute Tribe for water rights. The technical reality vendors won't tell you: Vector database costs don't scale linearly. That $5K/month at 2TB becomes $75K/month at 10TB. RAG inflates token usage from 15 to 500+ per query. LLM inference costs dominate 60-80% of total expenses. Semantic chunking costs 37.5% more but delivers 15-20% better accuracy, the difference between success and user complaints. The question isn't whether RAG works technically. It clearly does. The question is: who pays for "cheaper" AI? Right now, the answer is Kenyan workers at $1.32/hour. Publishers without licensing deals. Communities bearing environmental costs they didn't create. Enterprises discovering that 72% failure rates make "cheaper" just deferred expensive. This isn't inevitable. It's a choice about who benefits from human intelligence. I documented everything: the labor chains, the copyright battles, the environmental data, the enterprise failures, the infrastructure consolidation, and the cooperative alternatives that exist in proof-of-concept form right now. Six months. 50+ primary sources. Court documents. Financial reports. Worker testimony. Technical analysis. The full investigation reveals the voices Silicon Valley doesn't want you to hear, and the alternative futures they're fighting to create. 🔗 Read the complete investigation on Aylgorith: https://lnkd.in/d2FCtMx4 #AIEconomics #DigitalColonialism #TechAccountability
-
Estimated costs to train bio-sequence models, based on their published methods and cloud-compute pricing: 𝑃𝑟𝑜𝑠𝑡𝑇5 • 'finetuned for 10 days on 8 NVIDIA A100 each with 80GB vRAM' • $2,359 min (Spot instances) - $7,864 max (On-demand) 𝑃𝑟𝑜𝑡𝐺𝑃𝑇2 • 'trained on 128 NVIDIA A100s in 4 days' • $15,099 (min) - $50,335 (max) 𝑍𝑦𝑚𝐶𝑇𝑅𝐿 • '48 NVIDIA A100s 80GB for about 15,000 GPU hours' • $18,431 (min) - $61,444 (max) 𝐷𝑁𝐴𝐵𝐸𝑅𝑇-2 • '14 days using eight NVIDIA RTX 2080Ti GPUs' • $790 (min) - $2,628 (max) 𝐷𝑁𝐴𝐵𝐸𝑅𝑇 • '25 days on 8 NVIDIA 2080Ti GPUs' • $1,410 (min) - $4,692 (max) 𝐸𝑆𝑀-1𝑣 • 'ESM-1v models [...] for 6 days on 64 V100 GPUs. Weights for the MSA Transformer [...] 13 days on 128 V100 GPUs' • $57,508 (min) - $191,816 (max) 𝐺𝑒𝑛𝑆𝐿𝑀 • '40 A100 GPUs [...] approximately 6 hours' • $295 (min) - $983 (max) 𝑃𝑟𝑜𝑡𝑒𝑖𝑛𝐵𝐸𝑅𝑇 • 4 weeks on a single GPU • $155 (min) - $504 (max) Consider adding labor costs and GPU hours to develop the models on top of that. Breaking down the cost, using ProstT5 as an example: • GPU hours (# of GPUs × hours per GPU): 1,920 • Instance hourly cost: $9.83 (AWS Spot Instance) (8x A100) • Instance hourly cost: $32.77 (AWS On-Demand) (8x A100) • Number of instances: 1 • Hours per instance: 240 • Est. cloud cost: $2,359 (min, spot) - $7,864 (max, on-demand) Given that the cost per kWh here is about $0.35 and most of the necessary GPUs pull between 170W and 600W, it can be an order of magnitude cheaper to train with on-prem GPUs and only pay the cost of electricity. But what about the capital expense to purchase GPUs and build an on-prem server? Yes, even after that.
-
They tell you training and running AI model costs billions. That's true for a few frontier labs. But for most real-world use cases? Dramatically lower than you think thanks to open-source. Real examples from @HuggingFace's latest analysis: - Fine-tune a text classification model: <$2k - Train a leading image embedding model: < $7k - Train Deepseek OCR: < $100k - Train a leading machine translation model: <$500k Compare that to GPT-4.5 training (~$300M est.) And the truth is that you don't need a Formula 1 car to pick up groceries. Most tasks are solved just as well by smaller, efficient, targeted models. The mistake everyone makes? Starting with "what's the best AI model?" instead of "what do I need to do?" The future of AI is not just bigger models. It's cheaper, more customized, open models solving specific problems. Explore 100+ real model costs of training and deployment yourself in the study!
-
A new paper from Epoch AI shows that the cost of training frontier AI models has been growing at an exponential rate—2.4x per year since 2016. If this trend holds, by 2027, training the largest models will cost over $1 billion. The study breaks down the cost structure of training 45 leading AI models, including GPT-4 and Gemini Ultra, revealing that: —> 47–67% of costs go toward hardware (GPUs, TPUs, and servers) —> 29–49% of costs are R&D staff expenses, including salaries and equity —> Only 2–6% of costs go to energy consumption, but power demands are rising fast The financial barrier to training state-of-the-art models is becoming prohibitive, concentrating AI development in the hands of a few well-funded players and raises a few questions, especially after the recent release of the efficient DeepSeek model: Can new efficiency breakthroughs, alternative architectures, or shifts in funding models keep AI development open to more players? Full analysis https://lnkd.in/gSC_kkBU — Join thousands of world-class researchers and engineers from Google, Stanford, OpenAI, and Meta staying ahead on AI http://aitidbits.ai
-
Sure, it's widely known that pretraining large language models (LLMs) is incredibly expensive, but how expensive, exactly? Back in January, I did an back-of-the-envelope calculation: """ A 7B Llama 2 model costs about $760,000 to pretrain! And this assumes you get everything right from the get-go: no hyperparameter tuning, no debugging, no crashes and restarting. This also doesn't include personnel cost! The real cost, including experimentation, is probably millions! It makes me appreciate all these openly available models! Math: - The total number of GPU hours needed is 184,320 hours. - The cost of running one A100 instance per hour is approximately $33. - Each instance has 8 A100 GPUs. That's 184320 / 8 * 33 = $760,000 """ Now, the DeepSeek team now just released their latest v3 model, including their own cost estimate. Training the ~685B parameter model comes to about $5 million, assuming the sticker price of rented GPUs! Again, this doesn’t account for hyperparameter tuning costs, failed runs, or personnel expenses. (Of course, this is based on the sticker price of rented GPUs. And many large companies have in-house clusters or negotiate discount deals with cloud providers. But still...) In any case, my deepest appreciation goes out to the many researchers, engineers, and companies who share model weights openly so that we can to use, tinker, and post-train them! Link to the report on GitHub: https://lnkd.in/g8VKm3ZP
Explore categories
- Hospitality & Tourism
- Productivity
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Science
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Employee Experience
- Healthcare
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Career
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development