Production Inference: the Performance Angle [Partner talk] [ukr]

Deploying and Optimizing LLM Inference in Production on GPUs Ranging from L4 to 4xH200: Choosing Inference Engines and Configurations, Balancing Throughput, Latency, and Cost, and How to “Squeeze” the Most Out of Your Hardware and Avoid Common Tuning Mistakes

Dmytro Fedorenko
AI Director at De Novo
  • De Novo has currently built two cloud platforms for AI/ML: De Novo customers use LLMs, Speech To Text, CV;
  • Using artificial intelligence, Dmytro helps companies implement intelligent search, document analysis, and corporate knowledge management;
  • According to Dmytro, multimodal generative models, personalized generative systems, predictive analytics, and contextual understanding of natural language will be most in demand in the coming years.
Sign in
Or by mail
Sign in
Or by mail
Register with email
Register with email
Forgot password?