Michał Wojdylak

AI Infrastructure Engineer

Building production AI systems, LLM infrastructure, inference platforms and cloud-native ML solutions.

Latest writing

All posts →
8 min read

Hosting Multiple Models on NVIDIA Triton Server Across AWS, Azure, and GCP

By consolidating multiple ML models onto a single GPU instance using NVIDIA Triton Inference Server, engineering teams can achieve up to a 90% reduction in cloud infrastructure costs while maintaining sub-millisecond latencies.

tritongpuawsazuregcpmlopsinferencemulti-model