Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products

Susan Chang explains how Elastic transitioned from siloed, ad-hoc AI agent evaluations to a unified, production-grade framework. She discusses balancing LLM-as-a-judge with deterministic rules, bridging Python data science evals with TypeScript production code, and implementing deep tracing to catch regressions across complex RAG and cybersecurity workloads while preserving domain context. By…

Image: InfoQ

Coverage 1 publisher

  1. InfoQ

    Presentation: Building Reusable Evaluation Frameworks for Agentic AI Products

Articles stay on their publishers’ sites; each link opens the original.