Michał Wojdylak

AI Infrastructure Engineer

Building production AI systems, LLM infrastructure, inference platforms and cloud-native ML solutions.

Latest writing

All posts →
10 min read

LLM Inference Benchmarks: vLLM vs SGLang vs TensorRT-LLM on a Single L40S

A systematic comparison of three leading LLM inference engines — vLLM, SGLang, and TensorRT-LLM — benchmarked on Qwen3-32B across concurrency scaling, long-context, quantization, and MoE workloads on a single NVIDIA L40S GPU.

llminferencevllmsglangtensorrt-llmbenchmarksl40sqwen3fp8quantizationmoe
8 min read

Hosting Multiple Models on NVIDIA Triton Server Across AWS, Azure, and GCP

By consolidating multiple ML models onto a single GPU instance using NVIDIA Triton Inference Server, engineering teams can achieve up to a 90% reduction in cloud infrastructure costs while maintaining sub-millisecond latencies.

tritongpuawsazuregcpmlopsinferencemulti-model