Evaluating Project Performance Metrics

Explore top LinkedIn content from expert professionals.

  • View profile for Armand Ruiz
    Armand Ruiz Armand Ruiz is an Influencer

    building AI systems @meta

    207,232 followers

    Explaining the Evaluation method LLM-as-a-Judge (LLMaaJ). Token-based metrics like BLEU or ROUGE are still useful for structured tasks like translation or summarization. But for open-ended answers, RAG copilots, or complex enterprise prompts, they often miss the bigger picture. That’s where LLMaaJ changes the game. 𝗪𝗵𝗮𝘁 𝗶𝘀 𝗶𝘁? You use a powerful LLM as an evaluator, not a generator. It’s given: - The original question - The generated answer - And the retrieved context or gold answer 𝗧𝗵𝗲𝗻 𝗶𝘁 𝗮𝘀𝘀𝗲𝘀𝘀𝗲𝘀: ✅ Faithfulness to the source ✅ Factual accuracy ✅ Semantic alignment—even if phrased differently 𝗪𝗵𝘆 𝘁𝗵𝗶𝘀 𝗺𝗮𝘁𝘁𝗲𝗿𝘀: LLMaaJ captures what traditional metrics can’t. It understands paraphrasing. It flags hallucinations. It mirrors human judgment, which is critical when deploying GenAI systems in the enterprise. 𝗖𝗼𝗺𝗺𝗼𝗻 𝗟𝗟𝗠𝗮𝗮𝗝-𝗯𝗮𝘀𝗲𝗱 𝗺𝗲𝘁𝗿𝗶𝗰𝘀: - Answer correctness - Answer faithfulness - Coherence, tone, and even reasoning quality 📌 If you’re building enterprise-grade copilots or RAG workflows, LLMaaJ is how you scale QA beyond manual reviews. To put LLMaaJ into practice, check out EvalAssist; a new tool from IBM Research. It offers a web-based UI to streamline LLM evaluations: - Refine your criteria iteratively using Unitxt - Generate structured evaluations - Export as Jupyter notebooks to scale effortlessly A powerful way to bring LLM-as-a-Judge into your QA stack. - Get Started guide: https://lnkd.in/g4QP3-Ue - Demo Site: https://lnkd.in/gUSrV65s - Github Repo: https://lnkd.in/gPVEQRtv - Whitepapers: https://lnkd.in/gnHi6SeW

  • View profile for Bryan Howard

    I help CEOs, operators, and people leaders find and fix the drag slowing people, leaders, systems, and execution.

    29,846 followers

    "Why does our top performer get the worst reviews?" the VP asked me. I was reviewing their annual performance data. "Show me," I said. She pulled up the ratings. Diana: 2.8 out of 5. Below average on "collaboration." Low marks for "team player." "What's her actual performance?" I asked. "Exceeded every target. Landed our biggest client. Trained three new hires." "So why the low scores?" "Her peer reviews are dragging her down." I scanned the comments. "Too direct." "Challenges ideas too much." "Not supportive enough." "Let me talk to Diana," I said. "I used to give honest feedback," Diana told me. "Said our pricing model was broken. Got dinged for 'negativity.'" "What happened with the pricing?" "They finally fixed it six months later. After we lost two major accounts." "What else?" "I questioned why we needed  eleven approvals for a simple contract change. Manager said I wasn't being collaborative." "Are you still giving feedback?" "No. I learned my lesson. Now I smile. Nod. Say everything's great. My reviews are improving." "But nothing's actually improving?" "We're making the same mistakes. Just with better vibes." She chuckled. I went back to the VP. "Your review system doesn't measure performance," I said. "It measures compliance." "That's not true." "When was the last time someone got promoted for challenging bad ideas?" Silence. "When did someone get rewarded for preventing a mistake?" More silence. "You've trained your best people to stay quiet. And your mediocre people to stay nice." A few months later, they redesigned the system. Added a category: "Constructive Challenge." Points for identifying problems early. Rewards for preventing costly mistakes. Diana got promoted. "What changed?" I asked the VP. "We stopped confusing agreement with alignment. Stopped mistaking silence for harmony." "And?" "Turns out our 'difficult' people were our most valuable. They actually cared enough to speak up." Here's the truth about performance reviews: Most companies don't reward performance. They reward performance theater. The person who says the meeting was great beats the person who says it wasted an hour. The person who agrees with bad ideas beats the person who prevents disasters. You think you're measuring contribution. You're measuring conformity. And your best people? They've already figured out the game. They're just deciding whether to play it or find somewhere that values truth over comfort. _____ Like my content? Give me a follow. Want to see more of it? Click the 🔔 on my profile.

  • View profile for Sohrab Rahimi

    Director, AI/ML Lead @ Google

    24,249 followers

    Evaluating LLMs is hard. Evaluating agents is even harder. This is one of the most common challenges I see when teams move from using LLMs in isolation to deploying agents that act over time, use tools, interact with APIs, and coordinate across roles. These systems make a series of decisions, not just a single prediction. As a result, success or failure depends on more than whether the final answer is correct. Despite this, many teams still rely on basic task success metrics or manual reviews. Some build internal evaluation dashboards, but most of these efforts are narrowly scoped and miss the bigger picture. Observability tools exist, but they are not enough on their own. Google’s ADK telemetry provides traces of tool use and reasoning chains. LangSmith gives structured logging for LangChain-based workflows. Frameworks like CrewAI, AutoGen, and OpenAgents expose role-specific actions and memory updates. These are helpful for debugging, but they do not tell you how well the agent performed across dimensions like coordination, learning, or adaptability. Two recent research directions offer much-needed structure. One proposes breaking down agent evaluation into behavioral components like plan quality, adaptability, and inter-agent coordination. Another argues for longitudinal tracking, focusing on how agents evolve over time, whether they drift or stabilize, and whether they generalize or forget. If you are evaluating agents today, here are the most important criteria to measure: • 𝗧𝗮𝘀𝗸 𝘀𝘂𝗰𝗰𝗲𝘀𝘀: Did the agent complete the task, and was the outcome verifiable? • 𝗣𝗹𝗮𝗻 𝗾𝘂𝗮𝗹𝗶𝘁𝘆: Was the initial strategy reasonable and efficient? • 𝗔𝗱𝗮𝗽𝘁𝗮𝘁𝗶𝗼𝗻: Did the agent handle tool failures, retry intelligently, or escalate when needed? • 𝗠𝗲𝗺𝗼𝗿𝘆 𝘂𝘀𝗮𝗴𝗲: Was memory referenced meaningfully, or ignored? • 𝗖𝗼𝗼𝗿𝗱𝗶𝗻𝗮𝘁𝗶𝗼𝗻 (𝗳𝗼𝗿 𝗺𝘂𝗹𝘁𝗶-𝗮𝗴𝗲𝗻𝘁 𝘀𝘆𝘀𝘁𝗲𝗺𝘀): Did agents delegate, share information, and avoid redundancy? • 𝗦𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆 𝗼𝘃𝗲𝗿 𝘁𝗶𝗺𝗲: Did behavior remain consistent across runs or drift unpredictably? For adaptive agents or those in production, this becomes even more critical. Evaluation systems should be time-aware, tracking changes in behavior, error rates, and success patterns over time. Static accuracy alone will not explain why an agent performs well one day and fails the next. Structured evaluation is not just about dashboards. It is the foundation for improving agent design. Without clear signals, you cannot diagnose whether failure came from the LLM, the plan, the tool, or the orchestration logic. If your agents are planning, adapting, or coordinating across steps or roles, now is the time to move past simple correctness checks and build a robust, multi-dimensional evaluation framework. It is the only way to scale intelligent behavior with confidence.

  • View profile for Pierre Le Manh
    Pierre Le Manh Pierre Le Manh is an Influencer

    President and CEO, PMI

    87,576 followers

    Projects are investments. And project success is when the value delivered, as seen by all stakeholders, is worth the effort and expense. Last year, in the largest research study in PMI’s history, we reframed projects this way and redefined project success accordingly. This led us to a broader vision for the profession, a mindset to achieve project success beyond project management success: M.O.R.E. = Manage Perceptions, Own Project Success, Reassess Relentlessly, Expand Perspective. (see here: https://lnkd.in/eg56sT3K) Today, PMI is releasing the second wave of our global Project Success research. It shows with clarity that when project professionals step up to M.O.R.E., project success multiplies. This year’s report, “Step Up: Redefining the Path to Project Success with M.O.R.E.”, is based on insights from thousands of projects across 19 countries and close to 6,000 respondents – project professionals, sponsors, executives and customers. A few headlines that stand out for me: ➡️ Global project success is holding steady. 50% of projects are rated successful, 13% are outright failures, giving us a Global Net Project Success Score (NPSS) of 37. ➡️ A vision of success is non‑negotiable. When there is a clear vision of success, NPSS rises to 41, compared to –18 when vision is missing. ➡️ Measurement matters. When teams apply the full “measurement trifecta” – define success upfront, build a measurement system, and track outcomes throughout – NPSS jumps from 31 to 54. ➡️ M.O.R.E. multiplies success. When project professionals consistently practice all four elements of M.O.R.E., NPSS soars to 94, compared to 27 when none are applied. Yet only 7% say they always do all four. Our takeaway is that of course methods matter, but without the right mindset and agency, our profession would miss a huge opportunity to increase its impact. Stakeholders are telling us they expect project professionals to lead: to shape perceptions of value, to own outcomes, to reassess as conditions change, and to connect projects to broader organizational and societal goals. This is a call to action. If you are a project professional, this report is an invitation to STEP UP in the journey from being a task manager at the start of your career to becoming an enterprise value leader. And for us at PMI, it is an invitation to support you all along this journey, with our community‑vetted knowledge platform and resources, our career‑long learning and development opportunities, and our gold‑standard professional certifications. I hope you will read the report, discuss it with your teams, and use it to elevate your impact, on your projects, your organizations, and our world. 👉 Download “Step Up: Redefining the Path to Project Success with M.O.R.E.” here: https://lnkd.in/e-bGjxgD Project Management Institute #ProjectSuccess #MORE #ProjectManagement #Leadership #PMI

  • View profile for Aishwarya Srinivasan
    Aishwarya Srinivasan Aishwarya Srinivasan is an Influencer
    646,141 followers

    Most people evaluate LLMs by just benchmarks. But in production, the real question is- how well do they perform? When you’re running inference at scale, these are the 3 performance metrics that matter most: 1️⃣ Latency How fast does the model respond after receiving a prompt? There are two kinds to care about: → First-token latency: Time to start generating a response → End-to-end latency: Time to generate the full response Latency directly impacts UX for chat, speed for agentic workflows, and runtime cost for batch jobs. Even small delays add up fast at scale. 2️⃣ Context Window How much information can the model remember- both from the prompt and prior turns? This affects long-form summarization, RAG, and agent memory. Models range from: → GPT-3.5 / LLaMA 2: 4k–8k tokens → GPT-4 / Claude 2: 32k–200k tokens → GPT-OSS-120B: 131k tokens Larger context enables richer workflows but comes with tradeoffs: slower inference and higher compute cost. Use compression techniques like attention sink or sliding windows to get more out of your context window. 3️⃣ Throughput How many tokens or requests can the model handle per second? This is key when you’re serving thousands of requests or processing large document batches. Higher throughput = faster completion and lower cost. How to optimize based on your use case: → Real-time chat or tool use → prioritize low latency → Long documents or RAG → prioritize large context window → Agentic workflows → find a balance between latency and context → Async or high-volume processing → prioritize high throughput My 2 cents 🤌 → Choose in-region, lightweight models for lower latency → Use 32k+ context models only when necessary → Mix long-context models with fast first-token latency for agents → Optimize batch size and decoding strategy to maximize throughput Don’t just pick a model based on benchmarks. Pick the right tradeoffs for your workload. 〰️〰️〰️ Follow me (Aishwarya Srinivasan) for more AI insight and subscribe to my Substack to find more in-depth blogs and weekly updates in AI: https://lnkd.in/dpBNr6Jg

  • View profile for Lubomila J.
    Lubomila J. Lubomila J. is an Influencer

    Group CEO Diginex │ Plan A │ Greentech Alliance │ MIT Under 35 Innovator │ Capital 40 under 40 │ BMW Responsible Leader │ LinkedIn Top Voice

    170,336 followers

    Today is Earth Overshoot Day. Think of it like this: if Earth had a bank account of natural resources, we just spent our entire yearly allowance. And we still have 5 months left in 2025. Here’s what that actually means for business. We’re using trees faster than forests can grow back. We’re catching fish faster than they can reproduce. We’re using fresh water faster than rain and snow can refill our rivers and aquifers. It’s like running a business where you spend next year’s budget this year, every year. Right now, we’re consuming resources 1.7 times faster than nature can replace them. That’s not sustainable math and it’s creating real business problems today. Material costs have jumped 40% in just three years. Companies are losing an average of $182 million annually to supply disruptions. The World Economic Forum now lists resource scarcity as one of the biggest risks facing businesses this decade. Companies that make their resources last longer are seeing incredible returns. Examples exist across industries and we are lucky to with many of them. Many companies have figured out how to make their products last 20+ years instead of 10, and turned that into a $20 billion side business fixing and reselling equipment. Others cut their material costs in half by designing everything to be reused. The numbers speak for themselves. When companies invest in using resources more efficiently, they typically see returns of 15-25% higher than competitors. Most resource-saving projects pay for themselves within 18-36 months, then keep saving money year after year. Energy upgrades alone can return 20-30% on investment. Waste reduction programs often pay back 5-10 times what you put in during the first year. Even simple water conservation can pay for itself in 6-12 months. The real advantage isn’t just cost savings but it’s building a business that works when resources get scarce and expensive. Companies that stretch what they have are more predictable, need less capital, and can charge more when everyone else is scrambling for materials. #earthovershootday #sustainability #circulareconomy #businessstrategy #resourceefficiency #innovation #futureofbusiness #supplychainmanagement

  • View profile for Greg Coquillo

    AI Platform & Infrastructure Product Leader | Scaling GPU Clusters for Frontier Models | Microsoft Azure AI & HPC | Former AWS, Amazon | Startup Investor | I deploy the supercomputers that allow AI to scale

    233,890 followers

    "The LLM works great." Works great… according to what? That's the question most AI teams skip, and it's why so many models look brilliant in demos and fall apart in production. Testing an LLM isn't one thing. It's six, and using only one of them is how trust quietly breaks. Here are the 6 methods for testing LLM output quality 👇 🔹Human Evaluation - the gold standard for nuance, tone, subtle errors. Slow and costly, but irreplaceable. 🔹Automated Metrics - BLEU, ROUGE, BERTScore, perplexity. Fast and repeatable, weak on meaning. 🔹Adversarial & Red-Teaming - stress tests for jailbreaks, prompt injection, hallucinations. Critical before launch. 🔹LLM-as-a-Judge - a strong model grades outputs. Scales human-like judgment cheaply (watch for bias). 🔹Task-Specific Evaluation - custom datasets that mirror production. Measures real business value. 🔹Benchmark Testing - MMLU, HellaSwag, GSM8K, HumanEval. Comparable across models; may miss real-world tasks. The takeaway: no single method covers everything. Layer them. Save this if you build with LLMs. Which do you trust most? 👇

  • View profile for Christian Wattig

    Lead Instructor, Wharton FP&A Program | Corporate Trainer | Founder, Inside FP&A | On-site FP&A training at your offices (US & CA) and self-paced online learning

    123,120 followers

    Most FP&A teams spend hours on variance analysis and still miss the real problem. After 15+ years at P&G, Unilever, and Squarespace, I've watched skilled analysts calculate every variance to the penny and still walk into the leadership meeting unprepared for the question that actually matters. The issue is often that variance analysis gets treated as a reporting exercise when it should be an investigation. Here's the three-step approach I teach as a corporate FP&A trainer: 𝗦𝘁𝗲𝗽 1: 𝗧𝗵𝗲 𝗪𝗵𝗮𝘁 Identify what actually happened. Compare actuals to forecast at the right level of detail. Too granular, you drown in data. Too high-level, you miss what matters. 𝗦𝘁𝗲𝗽 2: 𝗧𝗵𝗲 𝗪𝗵𝘆 Most teams stop at "sales were down 10%." But why? Volume or price? New customers or retention? One product line or across the board? This is where the analysis usually breaks down. 𝗦𝘁𝗲𝗽 3: 𝗧𝗵𝗲 𝗦𝗼 𝗪𝗵𝗮𝘁 (this is most crucial!) Connect the variance to business impact. A 10% sales miss is fine if it's a timing issue. It's a crisis if a competitor is taking share. The ARCTIC framework (in the infographic below) is what I use to pressure-test the "So What" - which is where most variance analysis falls short. The teams I've seen do this well stop reporting variances and start using them as a forward-looking signal. Root cause becomes a forecast adjustment. Pattern becomes prevention. Which of the three steps trips up your team most often? -Christian Wattig 𝗣.𝗦. 𝗪𝗮𝗻𝘁 𝗺𝗼𝗿𝗲 𝗙𝗣&𝗔 𝗳𝗿𝗮𝗺𝗲𝘄𝗼𝗿𝗸𝘀? 👉 𝗝𝗼𝗶𝗻 𝗺𝘆 𝗻𝗲𝘅𝘁 𝗳𝗿𝗲𝗲 𝗹𝗶𝘃𝗲 𝘁𝗿𝗮𝗶𝗻𝗶𝗻𝗴 𝗵𝗲𝗿𝗲: https://lnkd.in/e9fEFjmK ________________________________________________ I'm the Director of the FP&A Certificate Program at Wharton Online and a former finance leader at P&G, Unilever, and Squarespace. I've trained 1,000+ professionals at companies like Google, Merck, and Lowe's. Here's how I can help: 🚀 Inside FP&A Academy My flagship online course for FP&A professionals who want to level up. 🤖 AI for FP&A A crash course on using AI to work faster and smarter in finance. 🏢 Corporate Training On-site workshops to upskill your finance team. 🔗 Go to InsideFPA[𝘥𝘰𝘵]com to learn more.

  • View profile for Dawid Hanak
    Dawid Hanak Dawid Hanak is an Influencer

    Professor advising industry & SMEs on evidence-based business cases for net zero and technology appraisals | TEA, LCA, Financial modelling | Low-Carbon, CCUS, Hydrogen Advisory | Helping academics publish & make impact

    61,391 followers

    The harsh truth: Without proper techno-economic assessment, your net zero technology or project can be *just* a science experiment. Here’s what you need to know. Performing a techno-economic assessment (TEA) from the early stage of technology or project development will support your decision making. It will provide you with key insights into costs, benefits, and feasibility. Here’s a quick breakdown of the key steps in a TEA: 1. Define Your Goal and Scope • What are you trying to achieve with this assessment? • Set clear objectives, boundaries, and a functional unit (e.g., cost per ton of CO₂ avoided). 2. Design Your Process and System Boundaries • Map out the process with flow diagrams and identify all key input/output streams. • Establish clear boundaries to understand what’s included in the analysis. 3. Gather Data for Inventory Analysis • Collect data on capital expenditures (CAPEX), operating costs (OPEX), energy use, and material inputs. • Address gaps and uncertainties in data collection. 4. Perform Economic Modeling • Break down costs into CAPEX, OPEX, and variable costs. • Use tools like discounted cash flow (DCF) analysis to calculate metrics like NPV and ROI. 5. Assess Key Performance Indicators (KPIs) • Focus on critical metrics such as: • Cost per ton of CO₂ avoided • Energy efficiency • Payback period 6. Run Sensitivity and Uncertainty Analysis • Identify the most significant cost drivers and test assumptions under different scenarios. • Identify and understand financial risks 7. Interpret and Present Results • Link findings to actionable recommendations for optimization or decision-making. • Communicate results in a way that resonates with stakeholders (e.g., policymakers, investors). Pro Tip: Combine TEA with life cycle assessment (LCA) to address both economic and environmental impacts for a holistic evaluation. 💡 Want to learn how to build and apply a TEA for your net zero project? I’ll be hosting regular 2-day training sessions throughout 2025 to provide hands-on guidance and tools to evaluate your projects confidently. The first cohort will be announced later today (as I’m screening through 250 applications!) #CarbonCapture #Research #Scientist #Sustainability #NetZero #ChemicalEngineering #Professor

  • View profile for Aurimas Griciūnas
    Aurimas Griciūnas Aurimas Griciūnas is an Influencer

    Founder @ SwirlAI • Ex-CPO @ neptune.ai (Acquired by OpenAI) • UpSkilling the Next Generation of AI Talent • Author of SwirlAI Newsletter • Public Speaker

    186,679 followers

    This is how you measure your AI system as an AI Engineer 👇 For regular software you would track metrics like uptime, error rate, p95 latency. However, they say little about whether the system is fast where users feel it, affordable at scale or correct. Here are the metrics we track when building LLM systems. It is useful to group them by the question they answer: 𝟭. 𝗜𝘀 𝗶𝘁 𝗳𝗮𝘀𝘁? (𝗟𝗮𝘁𝗲𝗻𝗰𝘆) ➡️ Time to first token (TTFT): how long the user is exposed to a blank screen, the number that defines perceived latency. ➡️ Inter-token latency (ITL): how smoothly tokens stream after the first one. ➡️ End-to-end latency at p50 / p95 / p99, dominated by output length, track it per use case rather than globally. 𝟮. 𝗖𝗮𝗻 𝗶𝘁 𝘀𝗰𝗮𝗹𝗲? (𝗧𝗵𝗿𝗼𝘂𝗴𝗵𝗽𝘂𝘁 𝗮𝗻𝗱 𝗰𝗼𝘀𝘁) ➡️ Tokens per second per user vs total system throughput, the two trade off against each other on the same hardware. ➡️ Input and output tokens per request to measure your unit economics. ➡️ Cache hit rate - prompt caching is often the technique that reduces cost the most. ➡️ Cost per successful task, not cost per request, a cheap request that fails is a waste. 𝟯. 𝗜𝘀 𝗶𝘁 𝗰𝗼𝗿𝗿𝗲𝗰𝘁? (𝗤𝘂𝗮𝗹𝗶𝘁𝘆) ➡️ Task success rate on a labeled eval set, re-run on every prompt or model change. ➡️ Groundedness for RAG - is the answer supported by the retrieved context. ➡️ Retrieval precision@k and recall@k - generation cannot fix what retrieval never surfaced. ➡️ LLM-as-judge scores over time, calibrated against human labels. ➡️ User feedback signals: thumbs, edits to generated output, free form feedback. 𝟰. 𝗗𝗼𝗲𝘀 𝗶𝘁 𝗵𝗼𝗹𝗱 𝘂𝗽? (𝗥𝗲𝗹𝗶𝗮𝗯𝗶𝗹𝗶𝘁𝘆) ➡️ Error, timeout and rate-limit rates per provider. ➡️ Retry and fallback rate - how often you silently switch to a backup model. ➡️ Guardrail trigger and refusal rates. 𝟱. 𝗛𝗼𝘄 𝗱𝗼𝗲𝘀 𝘆𝗼𝘂𝗿 𝗮𝗴𝗲𝗻𝘁 𝗯𝗲𝗵𝗮𝘃𝗲? (𝗔𝗴𝗲𝗻𝘁 𝗺𝗲𝘁𝗿𝗶𝗰𝘀) ➡️ Tool-call error rate. ➡️ Steps and tokens per completed task - drift here means cost is rising while accuracy remains the same ➡️ Context window utilization - the early warning for compaction and truncation issues. Read more about this in my newsletter: https://lnkd.in/dhiscYbm ❗️ Latency and reliability show up on day one because standard infra emits them. Quality, cost per task, and agent behavior need deliberate instrumentation, and they are where AI systems fail in production. Which metric caught a real problem for you that the standard dashboards missed? 👇

Explore categories