The cost of a single AI run tells you very little about whether an automated workflow is economically viable. A run is only an attempt. It is not necessarily a completed piece of work. A better measure is: 𝐓𝐨𝐭𝐚𝐥 𝐰𝐨𝐫𝐤𝐟𝐥𝐨𝐰 𝐜𝐨𝐬𝐭 ÷ 𝐀𝐜𝐜𝐞𝐩𝐭𝐚𝐛𝐥𝐞 𝐜𝐚𝐬𝐞𝐬 𝐜𝐨𝐦𝐩𝐥𝐞𝐭𝐞𝐝 Both figures should cover the same period. The total cost should include more than model usage. It may also include: Retries and failed runs Tool and infrastructure costs Human review Approval delays Corrections and rework Monitoring and maintenance The expected cost of errors that escape review A workflow might cost 20p per run. That sounds inexpensive. But if it usually needs three attempts, requires human review and occasionally creates costly downstream work, the run cost is only a small part of the real cost. “Acceptable” also needs to be defined before the workflow is assessed. It might mean that the work: meets the required quality standard is completed within the agreed time needs no unplanned correction does not create additional work elsewhere This also affects how autonomy should be designed. Stable, repeatable steps may be better handled with conventional automation. Tasks involving language, uncertainty or exceptions may justify the use of a model. Human judgement should be added where the likely consequence of a mistake justifies the interruption. Human review is not automatically safer. It can create delays, fatigue and rubber-stamping. Review should be used where the risk it removes is greater than the cost and friction it introduces. Before approving or scaling an AI workflow, decision-makers should examine: cost per acceptable completed case median and 95th-percentile completion time retry and failure rates human review time escalation rate correction and rework costs frequency and severity of errors that escape review The average can make a workflow look efficient. The 95th percentile often reveals the smaller number of cases that consume most of the time, money and attention. The useful question is not: “How cheap is one AI run?” It is: “What does it cost to complete the work properly, including the cases that go wrong?”
Evaluating Workflows for Efficiency
Explore top LinkedIn content from expert professionals.
-
-
HR doesn’t need more dashboards. It needs better listening. Most people teams measure what’s easy…like engagement scores or turnover. But the best teams? They build feedback loops that help them predict problems, not just react to them. This post gives you 11 of the most useful, often-overlooked loops you can implement across the employee lifecycle: 🟢 Week 2 new hire check-ins (capture early impressions) 🟠 Post-interview surveys (from both sides) 🔵 Onboarding reviews (day 90 is your goldmine) 🟡 Skip-level 1:1s (cross-level truth-telling) 🟣 Quarterly team health check-ins (lightweight, manager-led) …and 7 more. 📌 Save this if: • You’re building a modern HR function • You want fewer “We should’ve seen this coming” moments • You believe listening is strategy Which feedback loop is missing in your company?
-
How far are we from having competent AI co-workers that can perform tasks as varied as software development, project management, administration, and data science? In our new paper, we introduce TheAgentCompany, a benchmark for AI agents on consequential real-world tasks. Why is this benchmark important? Right now it is unclear how effective AI is at accelerating or automating real-world work. We hear statements like: > AI is overhyped, doesn’t reason, and doesn’t generalize to new tasks > AGI will automate all human work in the next few years This question has implications for: - Companies: to understand where to incorporate AI in workflows - Workers: to get a grounded sense of what AI can and cannot do - Policymakers: to understand effects of AI on the labor market How can we begin on it? In TheAgentCompany, we created a simulated software company with tasks inspired by real-world work. We created baseline agents, and evaluated their ability to solve these tasks. This benchmark is first of its kind with respect to versatility, practicality, and realism of tasks. TheAgentCompany features four internal web sites: - GitLab: for storing source code (like GitHub) - Plane: for doing task management (like Jira) - OwnCloud: for storing company docs (like Google Drive) - RocketChat: for chatting with co-workers (like Slack) Based on these sites, we created 175 tasks in the domains of: - Administration - Data science - Software development - Human resources - Project management - Finance We implemented a baseline agent that can web browse and write/execute code to solve these tasks. This was implemented using the open-source OpenHands framework for full reproducibility (https://lnkd.in/g4VhSi9a). Based on this agent, we evaluated many LMs, Claude, Gemini, GPT-4o, Nova, Llama, and Qwen. We evaluated both success metrics and cost. Results are striking: the most successful agent w/ Claude was able to successfully solve 24% of the diverse real-world tasks that it was tasked with. Gemini-2.0-flash is strong at a competitive price point, and the open llama-3.3-70b model is remarkably competent. This paints a nuanced picture of the role of current AI agents in task automation. - Yes, they are powerful, and can perform 24% tasks similar to those in real-world work - No, they can not yet solve all tasks or replace any jobs entirely Further, there are many caveats to our evaluation: - This is all on simulated data - We focused on concrete, easily evaluable tasks - We focused only on tasks from one corner of the digital economy If TheAgentCompany interests you, please: - Read the paper: https://lnkd.in/gyQE-xZG - Visit the site to see the leaderboard or run your own eval: https://lnkd.in/gtBcmq87 And huge thanks to Fangzheng (Frank) Xu, Yufan S., and Boxuan Li for leading the project, and the many many co-authors for their tireless efforts over many months to make this happen.
-
Most change initiatives don't fail because of the change that's happening, they fail because of how the change is communicated. I've watched brilliant restructurings collapse and transformative acquisitions unravel… Not because the plan was flawed, but because leaders were more focused on explaining the "what" and "why" than on how they were addressing the fears and concerns of the people on their team. People don't resist change because they don't understand it. They resist because they haven't been given a compelling story about their role in it. This is where the Venture Scape framework becomes invaluable. The framework maps your team's journey through five distinct stages of change: The Dream - When you envision something better and need to spark belief The Leap - When you commit to action and need to build confidence The Fight - When you face resistance and need to inspire bravery The Climb - When progress feels slow and you need to fuel endurance The Arrival - When you achieve success and need to honor the journey The key is knowing exactly where your team is in this journey and tailoring your communication accordingly. If you're announcing a merger during the Leap stage, don't deliver a message about endurance. Your team needs a moment of commitment–stories and symbols that anchor them in the decision and clarify the values that remain unchanged. You can’t know where your team is on this spectrum without talking to them. Don’t just guess. Have real conversations. Listen to their specific concerns. Then craft messages that speak directly to those fears while calling on their courage. Your job isn't just to announce change, but to walk beside your team and help your team understand what role they play in the story at each stage. #LeadershipCommunication #Illuminate
-
Much of the discussion around AI focuses on adoption rates, model capabilities, or productivity gains. The more consequential shift may be how organizations redesign work around it. Most organizations began their AI journey by applying new tools to existing tasks. Drafting content, conducting research, analyzing information, and automating routine tasks are often among the first use cases. Those applications are valuable, but they represent only the beginning. The larger opportunity emerges when organizations start redesigning workflows around AI rather than simply inserting AI into existing processes. That affects more than efficiency. It changes how decisions are made, how expertise is distributed, how teams collaborate, and how value is created across the organization. Organizations that treat AI as a workflow transformation capability rather than a productivity tool are likely to realize far greater long-term value.
-
Harvard report: 71% of meetings are unproductive. 65% keep people from doing real work. The best leaders don’t run more meetings. They run the ones that matter. The goal: Keep teams aligned, focused, and moving fast. Without wasting time. Here’s how to run the meetings that actually move the business forward: Weekly 1:1 ↳ Let them drive the agenda ↳ Listen first, coach second ↳ End with 2–3 clear next steps Leadership Team Meeting ↳ Review key metrics in 10 minutes ↳ Focus on 2–3 high-impact issues ↳ End with clear decisions and owners Weekly Operating Review ↳ Review KPIs by exception ↳ Flag risks in revenue, churn, or ops ↳ Assign fixes with owners and due dates Quarterly Planning Session ↳ Review last quarter’s goals ↳ Debate and choose top 3 priorities ↳ Assign clear owners and resources Voice-of-Customer Session ↳ Bring 3 real pain points ↳ Let customers talk 70% of the time ↳ Follow up within 30 days Board or Investor Update ↳ Share the hard stuff first ↳ Highlight 1–2 metrics that matter ↳ Ask for help with specific challenges All-Hands ↳ Explain the why, not just the what ↳ Take live, unscripted questions ↳ End with one clear message You may not need all of these. Some might add a daily standup. But chances are, your company doesn’t need half the meetings on the calendar now. Use this list to audit what’s working. Cut what’s not. Your team will thank you for their time back. Better meetings = faster decisions, sharper focus, and real momentum. P.S. Does your company have too many meetings or just right? ♻️ Repost to help a leader in your network. Follow Eric Partaker for more productivity insights. — 📌 Want PDFs of this and 100+ free leadership resources? Get them here: https://lnkd.in/ekhxjakK
-
There’s a huge difference between ‘I got AI to do this amazing thing for social media points’ and ‘I got AI to do this thing that generates a lot of revenue for my business or our clients.’ Real-world AI is very different. Most agents require small language models. Large context windows and multiple rounds of model calls turn the unit economics of foundational models negative for many use cases. Everything we build for clients starts with local AI. We spend no more than 2 days trying to get the workflow running on the Dell Pro Max T2 in my office. If it won’t run locally, using a frontier model rarely changes that. We scale the agent to support a small set of early adopters. This phase is critical. An early adopter cohort has been trained to use agents at their earliest maturity phase. Most users would reject the agent in this raw form. But this phase is intended to rapidly improve the agent’s workflow integration, orchestration, and reliability. Human feedback from trained early adopters improves agent performance faster than any other approach I have found. We iterate on more than just the LLMs. This phase fills in the knowledge graph, improves tool usage, adds guardrails, and informs the usage of more traditional machine learning models to augment the agent. When improvements plateau, we assess the agent. It is only promoted if its impact on outcomes meets user or customer expectations. Is it valuable? How does it reorchestrate workflows? Can the business monetize it? We roll the agent out to an alpha release cohort to scale the feedback flywheel. At this point, we know we have something valuable. We’re trying to improve its reliability and handle more workflow variations before a wider launch. We only evaluate frontier model usage at this phase. We finally know enough to make targeted decisions about where in the workflow frontier model performance could make a big enough difference to be worth considering. The alpha release also reveals adoption barriers for the agent and reorchestrated workflow. Most agents require us to craft an adoption journey for users and customers. That typically includes training for internal users and a phased rollout for customers. When improvement plateaus again, the agent is ready for general release. The process takes 2-3 months, and only about 30% of the workflows we try in my office end up going the distance. Data and information architecture make a huge difference. One client with a very mature knowledge graph is seeing a workflow success rate of over 50%. Small models perform significantly better for their use cases. #DellProMax
-
Getting the right feedback will transform your job as a PM. More scalability, better user engagement, and growth. But most PMs don’t know how to do it right. Here’s the Feedback Engine I’ve used to ship highly engaging products at unicorns & large organizations: — Right feedback can literally transform your product and company. At Apollo, we launched a contact enrichment feature. Feedback showed users loved its accuracy, but... They needed bulk processing. We shipped it and had a 40% increase in user engagement. Here’s how to get it right: — 𝗦𝘁𝗮𝗴𝗲 𝟭: 𝗖𝗼𝗹𝗹𝗲𝗰𝘁 𝗙𝗲𝗲𝗱𝗯𝗮𝗰𝗸 Most PMs get this wrong. They collect feedback randomly with no system or strategy. But remember: your output is only as good as your input. And if your input is messy, it will only lead you astray. Here’s how to collect feedback strategically: → Diversify your sources: customer interviews, support tickets, sales calls, social media & community forums, etc. → Be systematic: track feedback across channels consistently. → Close the loop: confirm your understanding with users to avoid misinterpretation. — 𝗦𝘁𝗮𝗴𝗲 𝟮: 𝗔𝗻𝗮𝗹𝘆𝘇𝗲 𝗜𝗻𝘀𝗶𝗴𝗵𝘁𝘀 Analyzing feedback is like building the foundation of a skyscraper. If it’s shaky, your decisions will crumble. So don’t rush through it. Dive deep to identify patterns that will guide your actions in the right direction. Here’s how: Aggregate feedback → pull data from all sources into one place. Spot themes → look for recurring pain points, feature requests, or frustrations. Quantify impact → how often does an issue occur? Map risks → classify issues by severity and potential business impact. — 𝗦𝘁𝗮𝗴𝗲 𝟯: 𝗔𝗰𝘁 𝗼𝗻 𝗖𝗵𝗮𝗻𝗴𝗲𝘀 Now comes the exciting part: turning insights into action. Execution here can make or break everything. Do it right, and you’ll ship features users love. Mess it up, and you’ll waste time, effort, and resources. Here’s how to execute effectively: Prioritize ruthlessly → focus on high-impact, low-effort changes first. Assign ownership → make sure every action has a responsible owner. Set validation loops → build mechanisms to test and validate changes. Stay agile → be ready to pivot if feedback reveals new priorities. — 𝗦𝘁𝗮𝗴𝗲 𝟰: 𝗠𝗲𝗮𝘀𝘂𝗿𝗲 𝗜𝗺𝗽𝗮𝗰𝘁 What can’t be measured, can’t be improved. If your metrics don’t move, something went wrong. Either the feedback was flawed, or your solution didn’t land. Here’s how to measure: → Set KPIs for success, like user engagement, adoption rates, or risk reduction. → Track metrics post-launch to catch issues early. → Iterate quickly and keep on improving on feedback. — In a nutshell... It creates a cycle that drives growth and reduces risk: → Collect feedback strategically. → Analyze it deeply for actionable insights. → Act on it with precision. → Measure its impact and iterate. — P.S. How do you collect and implement feedback?
-
Over the last year, I’ve seen many people fall into the same trap: They launch an AI-powered agent (chatbot, assistant, support tool, etc.)… But only track surface-level KPIs — like response time or number of users. That’s not enough. To create AI systems that actually deliver value, we need 𝗵𝗼𝗹𝗶𝘀𝘁𝗶𝗰, 𝗵𝘂𝗺𝗮𝗻-𝗰𝗲𝗻𝘁𝗿𝗶𝗰 𝗺𝗲𝘁𝗿𝗶𝗰𝘀 that reflect: • User trust • Task success • Business impact • Experience quality This infographic highlights 15 𝘦𝘴𝘴𝘦𝘯𝘵𝘪𝘢𝘭 dimensions to consider: ↳ 𝗥𝗲𝘀𝗽𝗼𝗻𝘀𝗲 𝗔𝗰𝗰𝘂𝗿𝗮𝗰𝘆 — Are your AI answers actually useful and correct? ↳ 𝗧𝗮𝘀𝗸 𝗖𝗼𝗺𝗽𝗹𝗲𝘁𝗶𝗼𝗻 𝗥𝗮𝘁𝗲 — Can the agent complete full workflows, not just answer trivia? ↳ 𝗟𝗮𝘁𝗲𝗻𝗰𝘆 — Response speed still matters, especially in production. ↳ 𝗨𝘀𝗲𝗿 𝗘𝗻𝗴𝗮𝗴𝗲𝗺𝗲𝗻𝘁 — How often are users returning or interacting meaningfully? ↳ 𝗦𝘂𝗰𝗰𝗲𝘀𝘀 𝗥𝗮𝘁𝗲 — Did the user achieve their goal? This is your north star. ↳ 𝗘𝗿𝗿𝗼𝗿 𝗥𝗮𝘁𝗲 — Irrelevant or wrong responses? That’s friction. ↳ 𝗦𝗲𝘀𝘀𝗶𝗼𝗻 𝗗𝘂𝗿𝗮𝘁𝗶𝗼𝗻 — Longer isn’t always better — it depends on the goal. ↳ 𝗨𝘀𝗲𝗿 𝗥𝗲𝘁𝗲𝗻𝘁𝗶𝗼𝗻 — Are users coming back 𝘢𝘧𝘵𝘦𝘳 the first experience? ↳ 𝗖𝗼𝘀𝘁 𝗽𝗲𝗿 𝗜𝗻𝘁𝗲𝗿𝗮𝗰𝘁𝗶𝗼𝗻 — Especially critical at scale. Budget-wise agents win. ↳ 𝗖𝗼𝗻𝘃𝗲𝗿𝘀𝗮𝘁𝗶𝗼𝗻 𝗗𝗲𝗽𝘁𝗵 — Can the agent handle follow-ups and multi-turn dialogue? ↳ 𝗨𝘀𝗲𝗿 𝗦𝗮𝘁𝗶𝘀𝗳𝗮𝗰𝘁𝗶𝗼𝗻 𝗦𝗰𝗼𝗿𝗲 — Feedback from actual users is gold. ↳ 𝗖𝗼𝗻𝘁𝗲𝘅𝘁𝘂𝗮𝗹 𝗨𝗻𝗱𝗲𝗿𝘀𝘁𝗮𝗻𝗱𝗶𝗻𝗴 — Can your AI 𝘳𝘦𝘮𝘦𝘮𝘣𝘦𝘳 𝘢𝘯𝘥 𝘳𝘦𝘧𝘦𝘳 to earlier inputs? ↳ 𝗦𝗰𝗮𝗹𝗮𝗯𝗶𝗹𝗶𝘁𝘆 — Can it handle volume 𝘸𝘪𝘵𝘩𝘰𝘶𝘵 degrading performance? ↳ 𝗞𝗻𝗼𝘄𝗹𝗲𝗱𝗴𝗲 𝗥𝗲𝘁𝗿𝗶𝗲𝘃𝗮𝗹 𝗘𝗳𝗳𝗶𝗰𝗶𝗲𝗻𝗰𝘆 — This is key for RAG-based agents. ↳ 𝗔𝗱𝗮𝗽𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗦𝗰𝗼𝗿𝗲 — Is your AI learning and improving over time? If you're building or managing AI agents — bookmark this. Whether it's a support bot, GenAI assistant, or a multi-agent system — these are the metrics that will shape real-world success. 𝗗𝗶𝗱 𝗜 𝗺𝗶𝘀𝘀 𝗮𝗻𝘆 𝗰𝗿𝗶𝘁𝗶𝗰𝗮𝗹 𝗼𝗻𝗲𝘀 𝘆𝗼𝘂 𝘂𝘀𝗲 𝗶𝗻 𝘆𝗼𝘂𝗿 𝗽𝗿𝗼𝗷𝗲𝗰𝘁𝘀? Let’s make this list even stronger — drop your thoughts 👇
-
Hyperautomation has emerged as a game-changer in the technological landscape, changing how businesses streamline operations, reduce costs, and enhance efficiency. By combining AI, ML, and robotic process automation (RPA), it transformed industries. Gone are the days when automation was limited to assembly lines or customer service bots. Hyperautomation transforms everything — from crunching financial data to streamlining inventory management — into a unified, efficient digital ecosystem. For instance: ▶️ In warehouses, IoT devices monitor inventory and trigger restocking before shelves go empty ▶️ Financial tools like RPA bots process invoices while AI forecasts cash flow trends ▶️ ML algorithms pinpoint supply chain inefficiencies and suggest actionable fixes The result? A seamless, real-time operational flow that saves time, money, and resources. Gartner projects that by 2026, 30% of enterprises will automate more than half of their network activities- up from under 10% in 2023. In finance, AI algorithms detect fraudulent transactions faster than human analysts, while RPA tools manage expenses and generate reports in seconds. Customer service chatbots powered by natural language processing (NLP) handle routine queries, leaving human agents free to focus on high-stakes issues. In manufacturing, predictive maintenance minimizes costly machine downtime by identifying potential issues before they arise. AI-powered quality control systems catch product defects that human eyes might miss, while workflow automation optimizes resource allocation. In the ever-complex supply chain, hyperautomation ensures real-time responsiveness. AI systems analyze traffic and weather to optimize delivery routes, while IoT devices keep stock levels in check. The result? Faster deliveries, fewer errors, and significant cost savings. While the potential of hyperautomation is undeniable, it raises questions about its impact on human labor. Repetitive, low-skill jobs are at the highest risk of being replaced. But, this shift also opens doors for workers to upskill to manage and optimize these systems, focusing on creative and strategic tasks instead of mundane ones. The narrative shouldn’t be “man versus machine” but “man with machine.” Valued at $45 billion in 2024, the hyperautomation market is projected to exceed $307 billion by 2037. Its future lies in driving sustainability, enabling hyper-personalized experiences, and achieving seamless end-to-end automation. As businesses continue to embrace this technology, it’s vital to maintain a human-centric approach: prioritizing ethical considerations, data privacy, and workforce training. The real question is: How will we harness its potential? #technology #AI #automation #innovation #business
Explore categories
- Hospitality & Tourism
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Science
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Artificial Intelligence
- Employee Experience
- Healthcare
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Career
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development