Gemma 4 on an Amazon SageMaker Endpoint: AWS CLI, NVIDIA L4, and an MCP Server

Chronological Source Flow
Back

AI Fusion Summary

This project demonstrates deploying Gemma 4 E2B to an Amazon SageMaker real-time endpoint using an NVIDIA L4 GPU and vLLM container. A suite of Python MCP tools simplifies management via Claude Code and AWS CLI. Performance tests show that the quantization-aware trained QAT checkpoint decodes 2.05x faster than the bf16 version, achieving 105.1 tok/s compared to 51.3 tok/s. Under 16 parallel requests, QAT serves 1077.25 tok/s against 619.1 tok/s while maintaining the same score.
Community Comments
Loading updates...
0