Evals for Testing AI Agents [ukr]

Building an AI agent is only half the job. The real challenge is knowing whether it performs reliably and continues to meet expectations as models, prompts, tools, and data evolve. In this session, we’ll use a practical example to show how to build effective Evals for AI agents: creating representative test cases, defining meaningful quality criteria and metrics, combining automated checks with LLM-as-a-Judge and human evaluation, comparing different agent versions, and making Evals a core part of the development and testing lifecycle.

Oleksander Krakovetskyi
СЕО at DevRain
  • CEO of DevRain, Ukrainian IT company
  • Co-founder and CTO of DonorUA - intellectual system for blood donors recruitment
  • PhD. in Computer Science
  • Microsoft Regional Director, Microsoft Artificial Intelligence Most Valuable Professional
  • Microsoft Certified: Azure Data Science Associate
  • Linkedin , Facebook
Sign in
Or by mail
Sign in
Or by mail
Register with email
Register with email
Forgot password?