Mercury: Blazing-Fast Generative AI on AWS
Inception Labs’ Mercury and Mercury Coder foundation models are now available on Amazon Bedrock Marketplace and SageMaker JumpStart, offering ultra-fast generative AI capabilities on AWS. Mercury models utilize a diffusion-based approach, generating multiple tokens in parallel for speeds up to 10x faster than comparable models (reaching 1,100 tokens/second on NVIDIA H100 GPUs). They excel at high-quality code generation across multiple languages (Python, Java, JavaScript, C++, PHP, Bash, TypeScript) and are particularly effective for code completion and editing. Both Bedrock Marketplace and SageMaker JumpStart provide streamlined deployment, scalable infrastructure, and robust security features. Bedrock offers a unified experience via its APIs and tools like Bedrock Agents and Guardrails, while SageMaker JumpStart integrates with SageMaker Studio and its MLOps features. Mercury models support context lengths up to 32,768 tokens (extensible to 128,000), and their transformer-based architecture ensures compatibility with existing optimization techniques. Deployment involves configuring endpoint details, instance types (e.g., ml.p5.48xlarge recommended for optimal performance), and security settings. The models can be tested in Bedrock’s playground before integration into applications. Examples showcase Mercury’s capabilities in generating a functional tic-tac-toe game and a travel planning assistant that interacts with external tools (weather API, calculator). Cleanup instructions are provided for both platforms to avoid unnecessary charges. The target audience includes developers and businesses seeking to integrate fast, high-quality generative AI into their applications.
Mercury leverages aws ai automation capabilities to deliver unprecedented speed and efficiency for enterprise generative AI deployments on Amazon Web Services.
While traditional chatgpt automation aws solutions often face latency issues, Mercury delivers unprecedented speed for enterprise-grade generative AI applications.

