Gemma-4-31B vLLM Serving PoC

EKS GPU01 benchmark topology - weight loading, serving path, and the L40S vs g7e verdict

Gemma-4-31B vLLM Serving PoC EKS GPU01 benchmark topology - weight loading, serving path, and the L40S vs g7e verdict EKS GPU01 - ap-northeast-2 (Seoul) vLLM pod - serving namespace Bench Client · vllm bench serve · Architecture component Bench Client vllm bench serve OpenAI API · vLLM /v1 endpoint · EKS GPU01 - ap-northeast-2 (Seoul) › vLLM pod - serving namespace OpenAI API vLLM /v1 endpoint vLLM Engine · v0.26.0 - Gemma-4-31B · EKS GPU01 - ap-northeast-2 (Seoul) › vLLM pod - serving namespace · serving ns vLLM Engine v0.26.0 - Gemma-4-31B serving ns MTP Draft Model · assistant - 56% accept · EKS GPU01 - ap-northeast-2 (Seoul) › vLLM pod - serving namespace MTP Draft Model assistant - 56% accept runai_streamer · weight loader · EKS GPU01 - ap-northeast-2 (Seoul) › vLLM pod - serving namespace runai_streamer weight loader S3 Weights · NVFP4 / FP8 checkpoints · Architecture component S3 Weights NVFP4 / FP8 checkpoints S3 Mountpoint · FUSE - legacy path · EKS GPU01 - ap-northeast-2 (Seoul) S3 Mountpoint FUSE - legacy path Karpenter · NodePool - RAID0 · EKS GPU01 - ap-northeast-2 (Seoul) Karpenter NodePool - RAID0 g6e L40S Nodes · 48GB - Marlin W4A16 · EKS GPU01 - ap-northeast-2 (Seoul) g6e L40S Nodes 48GB - Marlin W4A16 g7e Blackwell · 96GB - NvFp4 native, Spot · EKS GPU01 - ap-northeast-2 (Seoul) g7e Blackwell 96GB - NvFp4 native, Spot Compile Cache · hostPath VLLM_CACHE_ROOT · EKS GPU01 - ap-northeast-2 (Seoul) Compile Cache hostPath VLLM_CACHE_ROOT S1-S4 x c1-64 spec decode k=4 S3 direct 8.2 Gbps 18-32 s load FUSE 1.75 Gbps legacy 149 s compile hit 8.5 s Marlin fallback NvFp4 3.5-4.8x Legend Backend Database Cloud External

Throughput Verdict

  • • g7e beats L40S 3.49-4.83x on all 4 scenarios
  • • S1 chat: 536 -> 1,871 tok/s at c64
  • • KV cache 125,023 tokens removes preemption

Startup Reduction

  • • Cold 537 s -> warm 155 s on g6e.2xlarge
  • • runai_streamer S3 direct: 149 s -> 29 s weight load
  • • Persisted compile cache: 73 s -> 8.5 s

Korean Latency

  • • Official MTP assistant: 56% acceptance
  • • c1 E2E p50 6.85 s -> 2.69 s (2.6x)
  • • Full 262K vocab covers Korean, 0.94GB draft