Reviewing Progress Regularly

Explore top LinkedIn content from expert professionals.

  • View profile for Pan Wu
    Pan Wu Pan Wu is an Influencer

    Senior Data Science Manager at Meta

    52,000 followers

    A "sampled success metric" is a performance measure or evaluation criterion calculated from a sample or subset of data rather than the entire population. Its calculation often involves higher costs per sample, such as manual review, leading to a trade-off between sample size and metric accuracy/sensitivity. In this tech blog, written by the data science team from Shopify, the discussion revolves around how the team leverages Monte Carlo simulation to understand metric variability under various scenarios to help the team make the right trade-offs. Initially, the team defines simulation metrics to describe the variability of the sampled success metric. For instance, if the actual success metric is decreasing over time, the metric could indicate how many months of sampled success metric would show a decrease, termed as "1-month decreases observed". Then, the team defines the distribution to run the Monte Carlo simulation. Monte Carlo simulation, a computational technique using random sampling to estimate outcomes of complex systems or processes with uncertain inputs, draws samples from a dedicated distribution that matches business needs. Based on past observations, the team’s application follows a Poisson distribution. Next comes the massive simulation phase, where the team runs multiple simulations for one parameter and then changes various parameters to simulate different scenarios. The goal is to quantify how much the sample mean will differ from the underlying population mean given realistic assumptions. The final result provides a clear statistical distribution of how much extra sample size could lead to metrics variability decrease and increased accuracy. This case study demonstrates that Monte Carlo simulation could be a valuable toolkit to add to your decision-making and data science knowledge. #datascience #analytics #metrics #algorithms #simulation #montecarlo #decisionmaking – – –  Check out the "Snacks Weekly on Data Science" podcast and subscribe, where I explain in more detail the concepts discussed in this and future posts:    -- Spotify: https://lnkd.in/gKgaMvbh   -- Apple Podcast: https://lnkd.in/gj6aPBBY    -- Youtube: https://lnkd.in/gcwPeBmR https://lnkd.in/dKnrZzzV 

  • View profile for Niels Corsten

    Sr. Manager Service Design, CX & Journey Management @ Deloitte Digital

    5,685 followers

    A critical part of journey management in any large organisation is measuring how your journeys perform. 📊 By setting clear goals, monitoring performance, identifying gaps, and measuring improvement impact, you create a continuous cycle of management and enhancement. Measurement surfaces opportunities and kickstarts improvements. 🚀 Yet many organisations struggle: data sits in silos, teams measure inconsistently, and dashboards report numbers without a coherent story. Product, marketing, sales, service, and digital teams collect valuable insights, but without a common language, they never combine into a unified performance view. The result? Plenty of activity, little clarity on what actually improves customer experience and business performance. Measuring performance along specific journeys—rather than isolated KPIs—provides the right context: the journey itself. 🗺️ This approach transforms your journey framework into an engine for improving both customer experience and business performance holistically, creating a shared structure and language where different KPIs unite. 🧭 Inspired by the Balanced Scorecard, this pragmatic 3x3 Matrix structures performance measurement across two dimensions: 👉 First, it distinguishes 3 performance metric categories: - Customer performance (behavior and sentiment) - Commercial performance (conversion, customer base, revenue) - Operational performance (cost, efficiency, reliability) 👉 Second, it distinct three journey hierachy levels: - Overall customer lifecycle - End-to-end product or service journey - Individual customer tasks These intersecting dimensions ensure each metric sits logically within a complete, coherent view. The visual below shows example metrics for all nine sections, helping you build a balanced measurement framework for journeys. This matrix delivers three immediate benefits: ✨ 1. It aligns siloed KPIs and contextualizes them into a shared journey 2. It enables drill-down and aggregation through connected KPIs across journey levels 3. It surfaces trade-offs and synergies between performance metrics A few quick tips to take into account when drafting or structuring your own journey-driven measurement framework 👇👇👇 🐌 Consider both leading and lagging indicators for a robust measurement approach that balances early warning signs with outcome metrics.  🤲 Don’t collect everything. Start with a North Star KPI for each journey, and add a small set of supporting metrics. Less is more. 💬 Always mix performance metrics with more qualitative feedback and insights that will help you determine why performance is down and how to fix it. Happy measuring! 🎉

  • 𝗜𝗱𝗲𝗮 #𝟭𝟲: 𝗠𝗲𝘁𝗿𝗶𝗰𝘀 𝘁𝗵𝗮𝘁 𝗺𝗮𝘁𝘁𝗲𝗿: 𝘁𝗵𝗲 𝗯𝗲𝗮𝘂𝘁𝘆 𝗼𝗳 𝘀𝗽𝗶𝗹𝗹 𝗮𝗻𝗱 𝘀𝗽𝗼𝗶𝗹 I worked with a hotel chain that was focused on two high-level KPIs: 𝗮𝘃𝗲𝗿𝗮𝗴𝗲 𝗿𝗼𝗼𝗺 𝗿𝗮𝘁𝗲 (𝗔𝗥𝗥) and 𝗼𝗰𝗰𝘂𝗽𝗮𝗻𝗰𝘆 (%).  Occupancy was around 80% and had increased year on year but this aggregate average was hiding significant opportunities. When we de-averaged the overall occupancy by hotel and night, we discovered that very few hotels were 80% full: most were either completely full or only half full.  We reframed performance using two “failure metrics” (see illustration): • 𝗦𝗽𝗼𝗶𝗹: measured empty rooms (by hotel, by night). • 𝗦𝗽𝗶𝗹𝗹: measured “lost trading days” when a hotel reached full occupancy too early. By analysing 𝘀𝗽𝗶𝗹𝗹 𝗮𝗻𝗱 𝘀𝗽𝗼𝗶𝗹 𝗮𝘁 𝗮 𝘀𝗶𝘁𝗲-𝗻𝗶𝗴𝗵𝘁 𝗹𝗲𝘃𝗲𝗹, we uncovered significant value: • Spoil caused by pricing too high or insufficient marketing.   • Spill caused by pricing too low or overmarketing.   𝗦𝗽𝗼𝗶𝗹 𝗶𝘀 𝗮 𝗳𝗮𝗰𝘁. 𝗦𝗽𝗶𝗹𝗹 𝗶𝘀 𝗮 𝗺𝗼𝗱𝗲𝗹. One measures what you wasted; the other estimates what you missed.   The principle applies to almost any decision made under uncertainty: where there’s finite capacity and variable demand, there’s always a 𝘀𝗽𝗶𝗹𝗹-𝘀𝗽𝗼𝗶𝗹 𝘁𝗿𝗮𝗱𝗲-𝗼𝗳𝗳.  I’ve applied this framework across a diverse range of businesses: • 𝗖𝗮𝗹𝗹 𝗰𝗲𝗻𝘁𝗿𝗲𝘀: spill = calls with no agents (missed sales); spoil = agents with no calls (wasted labour). • 𝗥𝗲𝘀𝘁𝗮𝘂𝗿𝗮𝗻𝘁𝘀: spill = understaffed hours (poor service); spoil = overstaffed hours (low productivity). • 𝗦𝘂𝗽𝗲𝗿𝗺𝗮𝗿𝗸𝗲𝘁𝘀: spill = missed sales (poor availability); spoil = waste (over-stocking). Every business wrestles with these two-sided costs – the 𝗰𝗼𝘀𝘁 𝗼𝗳 𝗲𝘅𝗰𝗲𝘀𝘀 and the 𝗰𝗼𝘀𝘁 𝗼𝗳 𝗺𝗶𝘀𝘀𝗲𝗱 𝗼𝗽𝗽𝗼𝗿𝘁𝘂𝗻𝗶𝘁𝘆.  Once you measure both, you can manage the balance intelligently.  The best metrics don’t just describe performance – they expose 𝘧𝘢𝘪𝘭𝘶𝘳𝘦 𝘮𝘰𝘥𝘦𝘴 that can actually be fixed. Key takeaways: • Analyse at the most atomic level that could be actionable (hour, site-night, SKU-store, agent, keyword etc.) • Define the acceptable 𝗴𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹𝘀 for that atomic outcome.  • Systematically analyse the distribution of performance outside guardrails. • Recognise that averages hide opportunities where good and bad performance offset each other There’s a fascinating 140-year history of optimising these decisions which are commonly referred to as Newsvendor problems – but that story deserves its own post.

  • View profile for Johan Baltzar

    Co-Founder & CEO @ Steep | Hiring 🚀

    7,351 followers

    Metric trees – the power tool every data leader needs 🔧 So you’ve defined your company key metrics when one day a big metric drops – revenue is down. What’s going on? Your CEO wants to understand why – ASAP! Here’s where I found amazing value from using metric trees. Instead of that exercise spiraling into a confusing mess, with analysts and business folks looking into every metric, a few well-designed metric trees will help everyone focus and find the real underlying drivers. A basic example: • Your revenue is driven by two key levers: average order value and order volume. • Order volume depends on how many active users you have and how well they’re converting. • Conversion rate is impacted by landing page performance, page load speed, and traffic quality. When revenue dips, this structure gives you a starting point. Instead of poking around in dashboards, you can follow the path: Did order volume drop? Was traffic quality bad, or did conversion dip? Mathematically, there are no other ways for revenue to drop. Typically, you can find that one underlying driver is the culprit, allowing you to quickly rule out other hypotheses and focus efforts where it matters. It’s a shift from reactive analytics to proactive problem-solving. Is your org using metric trees?

  • View profile for Bahareh Jozranjbar, PhD

    UX Researcher at PUX Lab | Human-AI Interaction Researcher at UALR

    10,735 followers

    Benchmarking is one of the most direct ways to answer a question every UX team faces at some point: is the design meeting expectations or just looking good by chance? A benchmark might be an industry standard like a System Usability Scale score of 68 or higher, an internal performance target such as a 90 percent task completion rate, or the performance of a previous product version that you are trying to improve upon. The way you compare your data to that benchmark depends on the type of metric you have and the size of your sample. Getting that match right matters because the wrong method can give you either false confidence or unwarranted doubt. If your metric is binary such as pass or fail, yes or no, completed or not completed, and your sample size is small, you should be using an exact binomial test. This calculates the exact probability of seeing your result if the true rate was exactly equal to your benchmark, without relying on large-sample assumptions. For example, if seven out of eight users succeed at a task and your benchmark is 70 percent, the exact binomial test will tell you if that observed 87.5 percent is statistically above your target. When you have binary data with a large sample, you can switch to a z-test for proportions. This uses the normal distribution to compare your observed proportion to the benchmark, and it works well when you expect at least five successes and five failures. In practice, you might have 820 completions out of 1000 attempts and want to know if that 82 percent is higher than an 80 percent target. For continuous measures such as task times, SUS scores, or satisfaction ratings, the right approach is a one-sample t-test. This compares your sample mean to the benchmark mean while taking into account the variation in your data. For example, you might have a SUS score of 75 and want to see if it is significantly higher than the benchmark of 68. Some continuous measures, like task times, come with their own challenge. Time data are often right-skewed: most people finish quickly but a few take much longer, pulling the average up. If you run a t-test on the raw times, these extreme values can distort your conclusion. One fix is to log-transform the times, run the t-test on the transformed data, and then exponentiate the mean to get the geometric mean. This gives a more realistic “typical” time. Another fix is to use the median instead of the mean and compare it to the benchmark using a confidence interval for the median, which is robust to extreme outliers. There are also cases where you start with continuous data but really want to compare proportions. For example, you might collect ratings on a 5-point scale but your reporting goal is to know whether at least 75 percent of users agreed or strongly agreed with a statement. In this case, you set a cut-off score, recode the ratings into agree versus not agree, and then use an exact binomial or z-test for proportions.

  • View profile for Bill Carr

    Managing Partner - Board Trustee - Bestselling Author - Ex Vice President Amazon Video, Studios & Music

    24,922 followers

    One of the most common mistakes I see is when organizations rely on totals instead of rates when defining annual operating plan metrics. Many companies set goals like: “Deliver $461M in total cost improvement YoY.” At first glance, this sounds reasonable, but totals often obscure what’s really happening operationally. Absolute metrics are heavily influenced by volume, mix, and scale. If demand increases or decreases, the number moves—even if the underlying process hasn’t improved. For most operational metrics, a rate-based approach is far more useful. Rates isolate process performance from external factors. They normalize for scale, making it easier to determine whether the operation is actually improving. Instead of focusing on totals, define controllable input metrics such as: → Cost per unit shipped → Cost per delivery → Defects per thousand units → Delivery speed per hour Rates force clarity about what managers actually control. For example, a regional manager cannot fully control total delivery volume. But they can influence the cost per delivery, the defects per thousand units, or the delivery speed per hour. When metrics reflect what leaders actually control, ownership and accountability improve dramatically. During my time at Amazon, we would start each year with top-down output targets, such as revenue and fixed costs (especially headcount). But operational performance was largely managed through rate-based input metrics, including: a) Cost per unit shipped b) tp90 click-to-deliver time c) Demand-weighted in-stock rate d) Percent clicks in the top three search results e) tp90 page load time These metrics helped teams focus on sustainable improvements in operational excellence, regardless of demand fluctuations. Totals measure outcomes. Rates measure process performance. If you focus on process performance, desirable outcomes will follow.

  • View profile for Toby W.

    I help eCom brands scale past $25M/yr with Ads + Email Marketing. $450M+ in revenue | Moto, Leica, Kodak, Drake + 200+ more.

    23,298 followers

    Most brands analyze creative tests by looking at ROAS and CPA. That's like judging a restaurant by the bill instead of the food. ↳ Here's how to actually find winning patterns: Looking at performance metrics alone tells you IF something works. But it doesn't tell you WHY it works or how to replicate it. The Framework That Actually Works: 𝟭. 𝗦𝗽𝗹𝗶𝘁 𝗬𝗼𝘂𝗿 𝗠𝗲𝘁𝗿𝗶𝗰𝘀 𝗜𝗻𝘁𝗼 𝗧𝘄𝗼 𝗕𝘂𝗰𝗸𝗲𝘁𝘀 Primary metrics = Performance (tells you IF it works) - Spend, Purchases, CPA Secondary metrics = Storytelling (tells you WHY it works) - Scroll Stop Rate (hook strength) - Hold Rate (narrative engagement) - Outbound CTR (offer appeal) Why this matters: Performance metrics help you scale winners. Behavioral metrics help you create more winners. 𝟮. 𝗨𝘀𝗲 𝗕𝗲𝗵𝗮𝘃𝗶𝗼𝗿 𝘁𝗼 𝗙𝗶𝘅 𝗨𝗻𝗱𝗲𝗿𝗽𝗲𝗿𝗳𝗼𝗿𝗺𝗲𝗿𝘀 Don't change offers randomly. Let the data guide you: Low Scroll Stop Rate = Weak hook → Test bold claims, fast motion, pattern breaks Poor Hold Rate = Boring narrative → Improve pacing, cut slow parts Low Outbound CTR = Weak CTA/offer → Test different positioning Why this works: You're fixing the actual problem, not guessing at solutions. 𝟯. 𝗙𝗶𝗻𝗱 𝗣𝗮𝘁𝘁𝗲𝗿𝗻𝘀 𝗶𝗻 𝗬𝗼𝘂𝗿 𝗪𝗶𝗻𝗻𝗲𝗿𝘀 Stop looking at winning ads in isolation. Find common threads: Do they use specific hook styles? Similar pacing structures? Particular testimonial formats? Build a Creative Optimization Library documenting what works. Why this matters: Patterns create predictable processes. Processes eliminate guesswork. 𝟰. 𝗧𝗲𝘀𝘁 𝗪𝗶𝘁𝗵 𝗣𝘂𝗿𝗽𝗼𝘀𝗲 Most brands test random variations. Instead: If Scroll Stop Rate is bad → Test new hooks If Hold Rate is weak → Adjust storytelling If CTR is low → Optimize offer positioning Why this works: Every test has a clear objective and higher success probability. What You Can Expect: Fewer failed creative tests → Faster winner identification → Predictable creative production process → Higher overall ROAS from better optimization The Psychology: → Behavior data reveals true audience preferences. → Patterns show what actually drives action. → Purpose-driven testing eliminates waste. Next Steps: Week 1: Set up behavioral metric tracking Week 2: Analyze your last 10 winners for patterns Week 3: Build your Creative Optimization Library Week 4: Implement purpose-driven testing Be honest... Are you iterating creatives based on data, or gut instinct?

  • View profile for Daniel Svonava

    Self-host your inference, save $$$, own your AI | xYouTube

    40,409 followers

    Metrics Myopia: a common Information Retrieval affliction. 🧐📊 Symptoms include 95% precision but 0% user retention. Prescription: understand the metrics that actually matter. 💊 Order-Unaware Metrics: Precision in Simplicity 🎲 These metrics give you a straightforward view of your system's effectiveness, without worrying about results order. 1️⃣ Precision • What It Tells You: The accuracy of your retrieval—how many of the retrieved items are actually relevant. • When to Use: When users expect to get correct results right off the bat. 2️⃣ Recall • What It Tells You: The thoroughness of your retrieval—how many of all relevant items you managed to find. • When to Use: When missing information could be costly. 3️⃣ F1-Score • What It Tells You: The sweet spot between precision and recall, rolled into one metric. • When to Use: When you need to balance accuracy and completeness. Order-Aware Metrics: Ranking with Purpose 🏆 These metrics come into play when the order of results matters as much as the results themselves. 1️⃣ Average Precision (AP) • What It Tells You: How well you maintain precision across different recall levels, considering ranking. • When to Use: When assessing ranking quality for individual queries is crucial for your system's performance. 2️⃣ Mean Average Precision (MAP) • What It Tells You: Your system's average performance across multiple queries. • When to Use: For system evaluations, especially when comparing different models across diverse query types. 3️⃣ Normalized Discounted Cumulative Gain (NDCG) • What It Tells You: How well you're prioritizing the most relevant results and how quickly the first relevant result appears. • When to Use: In user-focused applications where top result quality can make or break the user experience. 4️⃣ Mean Reciprocal Rank (MRR) • What It Tells You: How quickly you're retrieving the first relevant item. • When to Use: When speed to the first correct answer is key, like in Q&A systems or chatbots. Choosing the Right Metric 🎯 The key is to align your metric choice with your system's goal. What matters most? • Precision? Go for Precision or MRR. • Completeness? Opt for Recall or F1-Score. • Ranking order? NDCG or MAP are your best bets. No single metric tells the whole story. Combine metrics strategically to gain a 360 review of your system's performance: • Pair Precision with Recall to understand both accuracy and coverage. • Use NDCG alongside MRR to evaluate both overall ranking quality and quick retrieval of top results. • Combine MAP with F1-Score to assess performance across multiple queries while balancing precision and recall. Finally, regularly reassess your metric choices as your system evolves and user needs change!

  • Is a metric tree a visual of metric relationships? A model of a business process? A possible analysis path? Or, a deep, rich analysis? It is all of the above rolled into one. Let’s take an example of a business process that moves from leads to contracts. When implemented as a metric tree, it not only shows how the individual metrics ladder up to the final output but it also encapsulates the underlying business process. But here is where it gets truly actionable. Inherently, this also represents an analysis decomposition path. If you are analyzing performance between two time periods, a product like HelloTrace can decompose down this metric tree and spotlight the specific metrics driving the drop in the output metric. What’s even more powerful is that you can take this “base metric tree” and further expand the decomposition to include segments like say marketing channels. The same analysis can now spotlight the specific segment drivers of the poor performance, and within each segment, the specific metric drivers contributing to the change. This is effectively a deep, rich analysis that is now automated and easily accessible. What this means is that the metric tree is far more than just a representation—it's a living, breathing analysis framework that models the business metrics, and enables detailed analysis in a way that is fast, deep, and operational. It’s the ultimate cheat code for data-driven decision-making.

  • View profile for Tony Christensen

    I help eCom brands scale past $25M/yr with Ads + Retention. $450M+ in revenue | Moto, Leica, Kodak, Drake + 200+ more.

    4,411 followers

    Scaling e-commerce brands isn’t about guesswork—it’s about data-driven creative strategy. Here’s the exact framework we use to turn creative insights into profit: 1. Shift From Guesswork to Data-Driven Decisions 🔹 Primary Metrics (Performance): Track purchases, CPA, and spend to know if your creative is working. 🔹 Secondary Metrics (Storytelling): Dive into Scroll Stop Rate, Hold Rate, Engagement Rate and Outbound CTR to understand why it’s working. Too many brands stop at CPA—without knowing why, you can’t replicate success. 2. Focus on Creative Optimization Build a Creative Optimization Feedback Loop to: ✅ Replicate winning elements across campaigns. ✅ Invest in high performers and cut underperformers. ✅ Refine creative weaknesses with precise insights. 3. Understand Consumer Behavior Primary metrics tell you the numbers; secondary metrics reveal consumer behavior. A low Hold Rate? It might be pacing or weak visuals. 4. Leverage Demographics & Placements Break down performance by age, gender, and placement to: 🔍 Discover hidden opportunities. 🎯 Personalize messaging for maximum impact. 💡 Tailor creatives for each segment. 5. Track Key Engagement Metrics Focus on: 👉 Scroll Stop Rate (grabs attention) 👉 Hold Rate (keeps attention) 👉 Engagement Rate (has emotion) 👉 Outbound CTR (drives traffic) Identify issues before they impact your budget. Final Thoughts: Systematic analysis > Guesswork. Data-driven creative wins. The best brands analyze, iterate, and scale—no gut feelings required. Action Plan: 1) Set up dashboards for performance and behavioral metrics. 2) Regularly review creative performance. 3) Use insights to refine and test new ideas. 4) Build a creative library of proven winners.

Explore categories