How do we bring AI to scientific modeling? The standard approach has been AI to augment existing numerical simulations. In a new work https://lnkd.in/gFMUvUbB we show this approach is fundamentally limited. In contrast, using the end-to-end AI approach of Neural Operators to completely replace numerical solvers helps overcome this limitation both in theory and in practice. Current augmentation approaches use AI as a closure model while keeping a coarse-grid numerical solver in the loop. We show that such approaches are generally unable to reach full fidelity, even if we make the closure models stochastic, providing them with history information and even unlimited ground-truth training data from full-fidelity solvers. This is because the closure model is forced to be at the same coarse resolution as the (cheap and approximate) numerical solver, and their combination does not result in high-fidelity solutions. In contrast, Neural Operators do not suffer from this limitation since they operate at any resolution and learn the mapping between functions. Neural Operators are first trained on coarse-grid approximate solvers, since we can generate lots of training data, and only use a small amount of expensive data from high-fidelity solvers in addition to physics-based losses to fine-tune the Neural Operator model for strong generalization. The key is that the Neural Operator model operates on any resolution, and can thus, accept data at multiple resolutions for training efficiently, without burdensome data-generation requirements. Thus, Neural Operators fundamentally change how we apply AI to scientific domains.
Overcoming Data Limitations In AI Model Development
Explore top LinkedIn content from expert professionals.
Summary
Overcoming data limitations in AI model development means finding ways to address the challenges posed by insufficient, incomplete, or low-quality data, which can hinder the performance and reliability of artificial intelligence systems. Solutions range from creating synthetic data to improving data quality and matching data types to specific AI tasks.
- Create synthetic data: Generate artificial datasets through simulation or modeling to fill gaps where real-world data is scarce or unavailable.
- Prioritize data quality: Invest time in cleaning, integrating, and validating your data to ensure models learn from accurate and reliable sources.
- Match data to AI goals: Make sure the data you gather fits the unique requirements of your AI solution, whether it’s raw signals for perception tasks or structured tables for analytics.
-
-
I am convinced that the next breakthrough in AI will not come from better algorithms. It will come from better data. Today, the most advanced neural architectures are available to everyone. The real differentiator is what you feed them. And in most critical environments, the data needed to train reliable AI models is scarce, incomplete, or simply does not exist: - You cannot train a detection system on a threat scenario that has never been observed. - You cannot expose a diagnostic model to rare pathologies that occur once in ten thousand clinical cases. - You cannot validate an autonomous function on a platform that is still being designed. This is the fundamental bottleneck. And this is where synthetic data generation becomes a game changer. At Scalian, we are developing simulation methods specifically designed to improve the performance of AI systems. We have built high-fidelity simulation capabilities that can generate synthetic data at scale, physically modelled, automatically annotated, and representative of variability and conditions that real-world data alone could never cover. However, we do not simply deliver data. Unlike many, we have also developed training methodologies that allow AI models to bridge the gap between synthetic and real-world inputs, ensuring that what models learn in simulation transfers reliably to operational conditions. As such, we have developed significant expertise in training AI using synthetic data. This enables us to support our clients throughout the entire AI development process. To stay ahead, we continuously feed our AI results back into our simulators, identifying which physical phenomena matter most for model performance, and refining our simulation fidelity accordingly. The impact is measurable. Detection rates that no observation-based approach has ever achieved. False positive rates divided. Models that generalise across scenarios no sensor has ever captured. I believe this new paradigm will enable us to build AI that is faster to develop, less expensive to train, and more trustworthy to deploy, especially in sectors such as Defence, Aerospace, Energy, or Healthcare.
-
A team spent 6 weeks debating GPT-4 vs Claude vs Gemini. They never shipped. Meanwhile, a competitor picked a model in a week and spent the rest of their time on data. They launched in a month. The difference? They understood something most teams miss. 𝐀𝐈 𝐏𝐞𝐫𝐟𝐨𝐫𝐦𝐚𝐧𝐜𝐞 = 𝐌𝐨𝐝𝐞𝐥 𝐂𝐚𝐩𝐚𝐛𝐢𝐥𝐢𝐭𝐲 × 𝐃𝐚𝐭𝐚 𝐐𝐮𝐚𝐥𝐢𝐭𝐲 A decent model (1.0x) with excellent data (1.5x) = 1.50 A superior model (1.1x) with poor data (0.5x) = 0.55 When models are commodities, the data multiplier dominates. Every time. 𝐃𝐢𝐟𝐟𝐞𝐫𝐞𝐧𝐭 𝐀𝐈 𝐭𝐲𝐩𝐞𝐬 𝐧𝐞𝐞𝐝 𝐝𝐢𝐟𝐟𝐞𝐫𝐞𝐧𝐭 𝐝𝐚𝐭𝐚: 𝐏𝐞𝐫𝐜𝐞𝐩𝐭𝐢𝐨𝐧 𝐀𝐈 → 𝐑𝐚𝐰 𝐃𝐚𝐭𝐚 → Computer vision, speech-to-text, document OCR → Needs unprocessed signals: images, audio, PDFs → Garbage images in = garbage detection out 𝐀𝐧𝐚𝐥𝐲𝐭𝐢𝐜𝐚𝐥 𝐀𝐈 → 𝐒𝐭𝐫𝐮𝐜𝐭𝐮𝐫𝐞𝐝 𝐃𝐚𝐭𝐚 → Forecasting, fraud detection, recommendations → Needs clean tables with resolved entities → Fragmented customer IDs = fragmented predictions 𝐆𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐯𝐞 𝐀𝐈 → 𝐂𝐨𝐧𝐭𝐞𝐱𝐭 𝐃𝐚𝐭𝐚 → Chatbots, content generation, summarization → Needs embedded knowledge AI can retrieve → Missing context = hallucinations 𝐀𝐠𝐞𝐧𝐭𝐢𝐜 𝐀𝐈 → 𝐀𝐜𝐭𝐢𝐨𝐧 𝐃𝐚𝐭𝐚 → AI that actually does things, not just answers → Needs tool specs, constraints, workflows → Stale API definitions = broken agents 𝐓𝐡𝐞 𝐛𝐨𝐭𝐭𝐨𝐦 𝐥𝐢𝐧𝐞: Stop asking "which model should we use?" Start asking "does our data match what this AI capability actually needs?" Models are commodities. Your data is the differentiator. What's the data gap holding back your AI project right now? #AI #DataEngineering #MachineLearning
-
𝟰𝟯% 𝗼𝗳 𝗔𝗜 𝗽𝗿𝗼𝗷𝗲𝗰𝘁𝘀 𝗳𝗮𝗶𝗹 𝗯𝗲𝗰𝗮𝘂𝘀𝗲 𝗼𝗳 𝗱𝗮𝘁𝗮 𝗾𝘂𝗮𝗹𝗶𝘁𝘆 Yet most organizations spend 80% on models and 20% on data. Your AI is only as smart as your data is clean. The pattern repeats across industries 👇 📊 𝗧𝗵𝗲 𝗗𝗮𝘁𝗮 𝗤𝘂𝗮𝗹𝗶𝘁𝘆 𝗖𝗿𝗶𝘀𝗶𝘀 Informatica's 2025 CDO survey found: ➜ 43% cite data quality as #1 obstacle to AI success ➜ 57% report data is NOT AI-ready ➜ Only 5% of organizations have comprehensive data governance 📉 𝗪𝗵𝗮𝘁 𝗕𝗮𝗱 𝗗𝗮𝘁𝗮 𝗟𝗼𝗼𝗸𝘀 𝗟𝗶𝗸𝗲 The data exists but: → Lives in 47 different systems with no integration → Uses inconsistent formats and definitions → Contains unknown biases that propagate through AI → Lacks lineage—nobody knows where it came from → Has quality issues discovered only after deployment Gartner predicts 30% of GenAI projects abandoned by end of 2025 due to poor data quality. 𝗧𝗵𝗲 𝗗𝗮𝘁𝗮 𝗘𝘅𝗰𝗲𝗹𝗹𝗲𝗻𝗰𝗲 𝗙𝗿𝗮𝗺𝗲𝘄𝗼𝗿𝗸 Organizations achieving production AI allocate 50-70% of timeline and budget to data readiness. Here's what they build: 1. 𝗖𝗼𝗺𝗽𝗿𝗲𝗵𝗲𝗻𝘀𝗶𝘃𝗲 𝗔𝘀𝘀𝗲𝘀𝘀𝗺𝗲𝗻𝘁 Completeness: Do you have sufficient volume? Accuracy: Is the data correct? Consistency: Do definitions match across systems? Timeliness: Is data current enough for decisions? Validity: Does data conform to business rules? 2. 𝗟𝗶𝗻𝗲𝗮𝗴𝗲 & 𝗣𝗿𝗼𝘃𝗲𝗻𝗮𝗻𝗰𝗲 For every data point: Where did it originate? How was it transformed? What systems touched it? When was it last validated? You can't trust AI you can't trace. 3. 𝗕𝗶𝗮𝘀 𝗗𝗲𝘁𝗲𝗰𝘁𝗶𝗼𝗻 & 𝗠𝗶𝘁𝗶𝗴𝗮𝘁𝗶𝗼𝗻 identify: Sample bias (unrepresentative training data) Historical bias (past discrimination baked in) Measurement bias (flawed data collection) Aggregation bias (combining incompatible data) Then engineer mitigation before deployment. 4. 𝗔𝗜 𝗚𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲 requires: Model-specific data requirements documentation Continuous data quality monitoring Automated drift detection Regular revalidation cycles 5. 𝗗𝗮𝘁𝗮 𝗣𝗿𝗲𝗽𝗮𝗿𝗮𝘁𝗶𝗼𝗻 𝗜𝗻𝗳𝗿𝗮𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲 Build platforms that enable: Extraction from source systems Normalization and transformation Quality dashboards with real-time monitoring Retention controls meeting compliance requirements API access for AI consumption Data readiness is NEVER "complete." It's continuous discipline requiring dedicated ownership. The Data Excellence Test: Ask yourself these questions: ✓ Can you trace any data point from source to consumption? ✓ Can you explain its quality metrics and bias profile? ✓ Do you have automated systems detecting data drift? ✓ Can you demonstrate data governance to regulators? ✓ Do you spend more on data infrastructure than AI models? If you answered "no" to any of these, you're building on quicksand. ♻️ Repost if you've seen AI fail due to data problems ➕ Follow for Pillar 4 tomorrow: Governance & Risk 💭 What percentage of your AI budget goes to data readiness?
-
In AI, more training data makes models worse, not better. A new discovery that flips everything we thought we knew about large language models: The "more is better" assumption just got shattered by a multi-university research team comparing two versions of the same LLM. They discovered that models trained on too much data develop "progressive sensitivity" - becoming so fragile they actually perform worse when fine-tuned for specific tasks. When researchers compared OLMo-1B models trained on 2.3T vs 3T tokens, the overtrained version showed 2-3% lower accuracy across standard benchmarks. Why? There's an "inflection point" where additional pre-training damages a model's adaptability. Beyond this threshold, the model's parameters become increasingly brittle - like a chess player who's drilled one strategy so extensively they can't adjust to new approaches. They validated this by adding noise to models at different training stages. Those trained beyond the inflection point crashed in performance with minimal perturbation. Less-trained models remained robust under similar conditions. This creates a fundamental dilemma for AI labs racing to build bigger models with more data. The implications are massive for how we approach AI development. Instead of brute force data accumulation, developers need to optimize the entire training pipeline. Finding the sweet spot between base capability and adaptability becomes crucial. For smaller models like OLMo-1B, this inflection point emerged around 2.5T tokens. But researchers acknowledge these effects may vary based on model size, architecture, and data quality. Not all experts agree - some research suggests intentional overtraining can reduce inference costs in specific scenarios. As we connect the dots between data science and systems engineering, it's becoming clear that AI development needs a more nuanced approach. The field must shift from "more data at all costs" to strategic, quality-focused training. This research doesn't just challenge assumptions - it completely reframes how we should approach scaling AI systems. Sometimes in AI, less truly is more.
-
𝗕𝗿𝗲𝗮𝗸𝗶𝗻𝗴 𝘁𝗵𝗲 𝗕𝗮𝘁𝗰𝗵 𝗘𝗳𝗳𝗲𝗰𝘁 𝗕𝗮𝗿𝗿𝗶𝗲𝗿: 𝗔 𝗡𝗲𝘄 𝗙𝗿𝗮𝗺𝗲𝘄𝗼𝗿𝗸 𝗳𝗼𝗿 𝗥𝗼𝗯𝘂𝘀𝘁 𝗣𝗮𝘁𝗵𝗼𝗹𝗼𝗴𝘆 𝗔𝗜 Foundation models in pathology face a persistent challenge: site-specific batch effects that limit their ability to generalize across hospitals. A new framework offers a practical solution. Research from Hai Cao Truong Nguyen et al. introduces fmMAP (Foundation Model-based Manifold Approximation Pipeline), a method designed to reduce site-bias batch effects in pathology foundation models without requiring full model retraining. The approach addresses a critical bottleneck: even state-of-the-art foundation models encode institutional signatures that can compromise downstream task performance. Why site bias matters: When pathology images come from different medical centers, they carry technical variations—differences in tissue preparation, staining protocols, scanner characteristics, and imaging workflows. Foundation models inadvertently learn these institutional fingerprints, which then propagate to every downstream application built on top of them. The result? Models that work well at one hospital but stumble at another. The fmMAP approach: - Works directly on foundation model representations rather than requiring access to raw training data or model retraining - Uses a maximum a posteriori framework to adjust embeddings and reduce site-specific variance while preserving biological signal - Demonstrated effectiveness across multiple downstream tasks, showing improved cross-site generalization - Provides a computationally efficient alternative to full model retraining or complex data preprocessing pipelines What makes this practical: Unlike methods that require modifying the foundation model training process or accessing proprietary datasets, fmMAP can be applied post-hoc to existing models. This makes it particularly valuable for institutions that rely on pre-trained foundation models but need to ensure reliable performance across their own diverse data sources. The broader implication: As pathology AI moves toward clinical deployment, we need tools that bridge the gap between research performance and real-world reliability. Methods like fmMAP represent an important step toward foundation models that truly generalize—not just within carefully curated datasets, but across the messy, heterogeneous reality of clinical practice. paper: https://lnkd.in/e-VcHMwp code: https://lnkd.in/eAxtv9wb #DigitalPathology #FoundationModels #MedicalAI #MachineLearning #ComputationalPathology #AIinHealthcare #BatchEffect #ModelRobustness #ClinicalAI — Subscribe to 𝘊𝘰𝘮𝘱𝘶𝘵𝘦𝘳 𝘝𝘪𝘴𝘪𝘰𝘯 𝘐𝘯𝘴𝘪𝘨𝘩𝘵𝘴 — weekly briefings on making vision AI work in the real world → https://lnkd.in/guekaSPf
-
A lot of conversations about AI systems focus on the model. Which LLM to pick. How to prompt it. Whether to fine-tune. What matters just as much is the data that flows in. And in most enterprises, that data is messy. CRM notes with shorthand and typos. Tax documents missing fields. Email threads where the key decision is buried three replies deep. The limits of the data become the limits of the model. An advisor-facing agent can’t resolve a client’s retirement question if the income data is incomplete. A compliance workflow can’t review communications effectively if half the emails aren’t categorized correctly. At Advisor360°, we spend as much time on data quality as on models. That means normalizing across systems, filling gaps, and attaching metadata that makes records machine-readable. It also means instrumenting agents to surface when they can’t find what they need—so missing data gets corrected, not ignored. The result is that our AI teammates don’t just run faster—they run on solid ground. The quality of insights, recommendations, and actions comes directly from the reliability of the underlying data. Models matter. But without disciplined work on data quality, the best LLM in the world will still stumble on the basics.
-
Why Do So Many AI POCs Die Young? A 15-Year Industry Pattern We Still Haven’t Fixed For nearly fifteen years, the corporate world has lived through a recurring paradox in artificial intelligence: AI proof-of-concepts almost always look stunningly successful, yet their real-world deployments collapse with startling consistency. Every CEO and CTO has, at some point, celebrated a POC that achieved 95–99% accuracy, only to watch the same model fall flat during rollout, lose credibility, and eventually be removed quietly from the product roadmap. This pattern has been so common and so predictable that it deserves serious, structural analysis. The uncomfortable truth is that AI POCs rarely fail because of weak algorithms; they fail because the data used in POCs is limited. A POC dataset is almost always limited in scope: it originates from a single site, a curated timeframe, or a constrained operational context. It includes just enough variables to demonstrate a concept, but never enough to represent the full distributional space across which the model is expected to generalize. In the controlled world of a POC, the model performs like a gifted student taking an exam for which it has already seen most of the questions. Once deployed, however, the model is suddenly confronted with demographic variations, operational biases, context shifts, previously unseen combinations of attributes, and temporal drifts for which the training data offers no guidance. What the industry has historically called “model failure” is, in fact, the deterministic consequence of building models on data that is too structurally thin to capture the complexity of real-world states. This phenomenon is better described as data blindness, the condition in which a model does not merely lack data—it lacks awareness of what kinds of data it lacks. Detecting this blindness is mathematically non-trivial. Traditional model metrics such as accuracy or AUC do not quantify what the model has never encountered. The missing knowledge lives in the combinatorial explosion of all possible attribute interactions—interactions that no finite dataset can fully represent. This is why, for decades, AI engineers have delivered high-performing POCs without realizing that the entire evaluation process was blind to the true dimensionality of the problem. Targeted data procurement offers a mature and scientifically grounded remedy. The solution is not to simply collect more data—an expensive and often directionless approach—but to identify precisely which segments of the data universe are missing. Causal modeling, TMLE, and generalizability analysis illuminate the structural gaps: under-represented populations, absent confounders, context-dependent variables, and site-specific differences that meaningfully influence outcomes. This transforms data acquisition from a brute-force operation into a strategic one.
-
Many teams overlook critical data issues and, in turn, waste precious time tweaking hyper-parameters and adjusting model architectures that don't address the root cause. Hidden problems within datasets are often the silent saboteurs, undermining model performance. To counter these inefficiencies, a systematic data-centric approach is needed. By systematically identifying quality issues, you can shift from guessing what's wrong with your data to taking informed, strategic actions. Creating a continuous feedback loop between your dataset and your model performance allows you to spend more time analyzing your data. This proactive approach helps detect and correct problems before they escalate into significant model failures. Here's a comprehensive four-step data quality feedback loop that you can adopt: Step One: Understand Your Model's Struggles Start by identifying where your model encounters challenges. Focus on hard samples in your dataset that consistently lead to errors. Step Two: Interpret Evaluation Results Analyze your evaluation results to discover patterns in errors and weaknesses in model performance. This step is vital for understanding where model improvement is most needed. Step Three: Identify Data Quality Issues Examine your data closely for quality issues such as labeling errors, class imbalances, and other biases influencing model performance. Step Four: Enhance Your Dataset Based on the insights gained from your exploration, begin cleaning, correcting, and enhancing your dataset. This improvement process is crucial for refining your model's accuracy and reliability. Further Learning: Dive Deeper into Data-Centric AI For those eager to delve deeper into this systematic approach, my Coursera course offers an opportunity to get hands-on with data-centric visual AI. You can audit the course for free and learn my process for building and curating better datasets. There's a link in the comments below—check it out and start transforming your data evaluation and improvement processes today. By adopting these steps and focusing on data quality, you can unlock your models' full potential and ensure they perform at their best. Remember, your model's power rests not just in its architecture but also in the quality of the data it learns from. #data #deeplearning #computervision #artificialintelligence
Explore categories
- Hospitality & Tourism
- Productivity
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Science
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Employee Experience
- Healthcare
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Career
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development