
Introducing explicit prompt caching for OpenAI GPT-5.6 models on Amazon Bedrock
Quick Answer
OpenAI's GPT-5.6 models, including Sol, Terra, and Luna, are now available on Amazon Bedrock, featuring explicit prompt caching for optimized performance and cost savings.
Quick Take
The caching mechanism allows users to save 90% on cached inputs, enhancing efficiency in agentic workflows. These models support various reasoning levels and structured outputs, catering to diverse application needs.
Key Points
- GPT-5.6 models include Sol for complex tasks, Terra for balanced workloads, and Luna for high-volume tasks.
- Explicit prompt caching allows 90% cost savings on reused inputs for 30 minutes.
- Models support various reasoning levels: none, low, medium, high, and xhigh.
- Streaming responses deliver output as typed events for real-time applications.
- Structured output can be enforced with strict JSON schema compliance.
DeepSignal Analysis
What happened
OpenAI's GPT-5.6 models, including Sol, Terra, and Luna, are now available on Amazon Bedrock. These models feature explicit prompt caching, allowing users to save 90% on cached inputs, which is particularly beneficial for agentic workflows. The models cater to different reasoning levels and structured outputs, enhancing their versatility for various applications.
Key evidence
- The GPT-5.6 models are available on Amazon Bedrock with pay-per-token pricing and AWS security controls.
- Explicit prompt caching allows users to cache portions of prompts for reuse, with cached input billed at a 90% discount.
- The models support various reasoning levels, including Sol for complex tasks, Terra for balanced workloads, and Luna for high-volume tasks.
Why it matters
The introduction of explicit prompt caching in GPT-5.6 models on Amazon Bedrock could significantly reduce operational costs for businesses utilizing AI. By allowing users to cache and reuse prompts, the models enhance efficiency in workflows that require repetitive tasks. This capability, combined with the tiered reasoning levels, positions these models as adaptable solutions for a range of applications, from coding to summarization.
Source Excerpt
OpenAI GPT-5. 6 Sol, Terra, and Luna are now generally available on Amazon Bedrock, along with explicit prompt caching that gives you precise control over which parts of your prompt are cached and reused. Learn how to get started, set up explicit caching, and migrate existing GPT workloads to reduce inference cost.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from AWS Machine Learning
See more →
Build an explainable next-best-product recommendation system for banking on AWS
AWS presents a deep learning-based Next-Best-Product recommendation system for banks, utilizing Amazon SageMaker and PyTorch to enhance customer product predictions. This architecture leverages a multi-tower neural network for improved accuracy and explainability, addressing the complexities of customer data in financial services.

