EKS GPU01 benchmark topology - weight loading, serving path, and the L40S vs g7e verdict
Guided views
Explore this system
Step through curated paths without changing the source diagram.
Beat
Next
ReadyChapter 01 / 01
Guided chapter
Diagram guideExplore this system
Inspecting compiled semantics
E ExportT ThemeS Style0 Reset+ Zoom in- Zoom outEsc Close
Find a node
⌕/
No matching nodes
Semantic passport
Verified source
Authored reach
Route probeChoose a start node
Pick two semantic nodes on the diagram
Choose the source, then the destination. Direction matters.
Semantic lensCompare system roles
Choose up to two semantic kinds. One reveals its real traffic; two compare only direct authored relationships.
Choose a kind to inspect its nodes and touching relationships.
Semantic radar
Building overview
Click nodeDrag to pan
Throughput Verdict
• g7e beats L40S 3.49-4.83x on all 4 scenarios
• S1 chat: 536 -> 1,871 tok/s at c64
• KV cache 125,023 tokens removes preemption
Startup Reduction
• Cold 537 s -> warm 155 s on g6e.2xlarge
• runai_streamer S3 direct: 149 s -> 29 s weight load
• Persisted compile cache: 73 s -> 8.5 s
Korean Latency
• Official MTP assistant: 56% acceptance
• c1 E2E p50 6.85 s -> 2.69 s (2.6x)
• Full 262K vocab covers Korean, 0.94GB draft
Architecture diagram • Built with Archify • Create yours ↗ • Hover to trace • R route • Click to focus • +/− zoom • M radar • [/] views • P play story • T theme • E export