Yuvraj Garg · AI Systems & Infrastructure Architect
Infrastructure Architect
San Francisco, CARemote OKFull-time$140k – $220k / year
About this role
Yuvraj Garg · AI Systems & Infrastructure Architect
ML Systems ArchitectBengaluru, India · Remote · Open to relocation
Hi, I'm Yuvraj Garg
AI Systems & GPU Infrastructure Engineer. Production ML, vLLM, KubeRay, LangGraph
I engineer production-scale AI infrastructure, low-latency GPU serving pipelines, and robust multi-agent orchestration systems.
I own the translation layer from research to codebases people pay for, driving model cost and cold starts down, debugging GPU memory limits when clusters break, and ensuring 99.9% reliability.
Download Resume View Impact Metrics
Proven Expertise:Styldod · 5y LeadMITx MicroMastersRed Hat Certified × 34 Products Shipped Solo
6.9×
Cold Start Reduction
FLUX.2-klein-9B GPU memory snapshots
18.6×
Cost Optimization
DocumentAI cluster vs AWS Bedrock
24×
LLM Boot Improvement
110s → 2-7s on Modal GPU layers
30+
Models In Production
Image, VLM, and agentic microservices
Business Impact & Achievements
Engineering metrics verified in production.
Instead of abstract scores, I benchmark real-world parameters: hosting costs, cold start seconds, parallel worker execution, and custom inference compilation.
Cost & Scale
~70% saved
Inhouse Model Hosting
Self-hosted 30+ models at Styldod rather than utilizing external APIs, handling millions of monthly inference queries at 1/3 the cost.
Styldod Core Infra
18.6× cheaper
vLLM on EKS Cluster
Built a self-hosted pipeline (EKS + KubeRay + vLLM) for DocumentAI, processing 50K docs/day at $0.025/doc vs $0.466 on AWS Bedrock.
AWS EKS / vLLM Deployment · 2026
40%+ saved
Serverless GPU Orchestration
Eliminated always-on GPU standby instances by migrating image pipelines to AWS EKS. Used scale-to-zero during idle hours to cut vendor costs by 40%+.
Vendor Migration · 2025
Speed & Optimization
48s → 7s
FLUX.2 Memory Snapshots
Achieved 6.9× faster startup for image generation models on L40S GPUs using serverless runtime memory checkpointing.
Modal GPU Snap
Related opportunities
Senior Agentic AI Engineer @ Eigen Labs
Eigen Labs
San Francisco, CA$180k – $280k
View →
ML Ops Engineer (EMEA Remote) @ Pragmatike
Pragmatike
San Francisco, CA$140k – $220k
View →
Machine Learning Engineer @ Sift
Sift
San Francisco, CA$140k – $220k
View →
Machine Learning Engineer @ Reducto
Reducto
San Francisco, CA$140k – $220k
View →