Gemma 4 31B now runs on AWS three different ways
Gemma 4 31B now has three routes onto AWS and Google — Bedrock, SageMaker, and Google Cloud — and each bills differently.
Google DeepMind's Gemma 4 31B now reaches AWS three different ways, each billed differently.
Google's own Cloud program covers it directly — credits apply to Gemini and Gemma with no extra setup. On AWS, Bedrock already lists Gemma 4 31B as a managed, pay-per-token model, and Activate credits cover third-party models via Bedrock the same way they cover any other model there.
What's new, as of 14 September, is a third option: SageMaker JumpStart now deploys Gemma 4 31B — including an NVFP4-quantized build AWS says is 2.5x faster and uses two-thirds less memory — onto an instance you pay for by the hour, not by the token. That only beats Bedrock's metered rate once inference volume is high enough to keep the instance busy.
Three routes, three different bills for the same model — pick by how much of it you will actually run.