Apply

Yuvraj Garg · AI Systems & Infrastructure Architect

Infrastructure Architect

San Francisco, CARemote OKFull-time$140k – $220k / year

About this role

Yuvraj Garg · AI Systems & Infrastructure Architect ML Systems ArchitectBengaluru, India · Remote · Open to relocation Hi, I'm Yuvraj Garg AI Systems & GPU Infrastructure Engineer. Production ML, vLLM, KubeRay, LangGraph I engineer production-scale AI infrastructure, low-latency GPU serving pipelines, and robust multi-agent orchestration systems. I own the translation layer from research to codebases people pay for, driving model cost and cold starts down, debugging GPU memory limits when clusters break, and ensuring 99.9% reliability. Download Resume View Impact Metrics Proven Expertise:Styldod · 5y LeadMITx MicroMastersRed Hat Certified × 34 Products Shipped Solo 6.9× Cold Start Reduction FLUX.2-klein-9B GPU memory snapshots 18.6× Cost Optimization DocumentAI cluster vs AWS Bedrock 24× LLM Boot Improvement 110s → 2-7s on Modal GPU layers 30+ Models In Production Image, VLM, and agentic microservices Business Impact & Achievements Engineering metrics verified in production. Instead of abstract scores, I benchmark real-world parameters: hosting costs, cold start seconds, parallel worker execution, and custom inference compilation. Cost & Scale ~70% saved Inhouse Model Hosting Self-hosted 30+ models at Styldod rather than utilizing external APIs, handling millions of monthly inference queries at 1/3 the cost. Styldod Core Infra 18.6× cheaper vLLM on EKS Cluster Built a self-hosted pipeline (EKS + KubeRay + vLLM) for DocumentAI, processing 50K docs/day at $0.025/doc vs $0.466 on AWS Bedrock. AWS EKS / vLLM Deployment · 2026 40%+ saved Serverless GPU Orchestration Eliminated always-on GPU standby instances by migrating image pipelines to AWS EKS. Used scale-to-zero during idle hours to cut vendor costs by 40%+. Vendor Migration · 2025 Speed & Optimization 48s → 7s FLUX.2 Memory Snapshots Achieved 6.9× faster startup for image generation models on L40S GPUs using serverless runtime memory checkpointing. Modal GPU Snap

Related opportunities