AWS Adds Moonshot AI's Kimi K3 Model to Amazon Bedrock
Amazon Web Services announced that Kimi K3, the latest large language model from Moonshot AI, is now available on Amazon Bedrock. The AWS Machine Learning Blog states that Kimi K3 is the first open model to reach 2.8 trillion parameters and supports a 1-million-token context window. AWS says the model also introduces prompt caching to reduce costs and latency, and offers enhanced security with zero data retention for inference requests.
Show evidence (4 verified excerpts)
According to Moonshot AI, Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters.
This is the source’s account; it does not establish independent confirmation.
It combines native vision capabilities with a 1-million-token context window and delivers an approximate 2.5x improvement in scaling efficiency over Kimi K2.
This is the source’s account; it does not establish independent confirmation.
Kimi K3 is the first open-weight model on Amazon Bedrock to support explicit prompt caching, helping you reduce latency and input costs when reusing context across model calls.
This is the source’s account; it does not establish independent confirmation.
Zero data retention is always enabled for inference requests, while zero operator access prevents even AWS operators from accessing your prompts and completions during inference.
This is the source’s account; it does not establish independent confirmation.
High Signal1 min read
Kimi K3 Model Features
Kimi K3 is described by Moonshot AI as its most capable model to date, and the first open model to reach 2.8 trillion parameters. It includes native vision capabilities, a 1-million-token context window, and claims a 2.5x efficiency improvement over Kimi K2. [1]
Show evidence (2 verified excerpts)
According to Moonshot AI, Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters.
This is the source’s account; it does not establish independent confirmation.
It combines native vision capabilities with a 1-million-token context window and delivers an approximate 2.5x improvement in scaling efficiency over Kimi K2.
This is the source’s account; it does not establish independent confirmation.
Deployment, Security, and Cost Controls
AWS highlights that Kimi K3 brings explicit prompt caching, which aims to lower latency and input costs when reusing context. The platform enforces zero data retention and zero operator access for greater data security. Organizations can adopt Kimi K3 without changing their security posture. [1]
Show evidence (3 verified excerpts)
Kimi K3 is the first open-weight model on Amazon Bedrock to support explicit prompt caching, helping you reduce latency and input costs when reusing context across model calls.
This is the source’s account; it does not establish independent confirmation.
Zero data retention is always enabled for inference requests, while zero operator access prevents even AWS operators from accessing your prompts and completions during inference.
TechCrunch AI reports that Jev, a new type of AI model from a ChatGPT inventor, is attracting attention from developers by offering a cheaper and faster approach to building software intelligence.
The Verge reports that court documents unsealed in the New York Times' lawsuit reveal OpenAI and Microsoft's internal warnings about their own AI training practices. The companies' own records described their web data scraping as a 'doom loop' that could harm the internet, calling it the 'largest theft of labor in human history' and undermining 'fair use.'
TechCrunch AI reports that a hallucination by an AI large language model almost led to a US military operation. A GovAI research scholar highlights the critical need for service members to recognize 'the uncertainty inherent to LLMs.'