Kimi K3 Designs AI Chip in 48 Hours as Moonshot Launches 2.8 Trillion Parameter Model
Chinese artificial intelligence company Moonshot AI has introduced Kimi K3, a 2.8 trillion parameter Mixture of Experts model designed for advanced coding, multimodal reasoning, knowledge work, and long duration autonomous tasks. Moonshot describes Kimi K3 as the first open model within the 3 trillion parameter class, although its complete model weights are scheduled to be released by July 27, 2026. The model is already available through Kimi, Kimi Work, Kimi Code, and the company’s API.
Kimi K3 uses Moonshot’s Kimi Delta Attention and Attention Residuals architectures, alongside Stable LatentMoE. The model contains 896 experts but activates only 16 for each operation, allowing it to access an enormous parameter base without using every parameter during each inference step. Moonshot claims these architectural and training improvements deliver approximately 2.5 times better scaling efficiency than Kimi K2.
Introducing Kimi K3: Open Frontier Intelligence
— Kimi.ai (@Kimi_Moonshot) July 16, 2026
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional… pic.twitter.com/eFHEbdxn3P
The primary technical specifications include:
Total parameters: 2.8 trillion
Architecture: Mixture of Experts with 16 of 896 experts activated
Context window: 1 million tokens
Multimodal support: Native text, image, and video understanding
Core technologies: Kimi Delta Attention, Attention Residuals, Gated MLA, and Stable LatentMoE
Data formats: MXFP4 weights and MXFP8 activations
Reasoning: Maximum reasoning effort enabled by default at launch
Availability: Kimi, Kimi Work, Kimi Code, and the Kimi API
Open weights: Scheduled for release by July 27, 2026
API pricing: $0.30 per 1 million cached input tokens, $3 per 1 million standard input tokens, and $15 per 1 million output tokens
Moonshot acknowledges that Kimi K3 still trails the strongest proprietary systems, including GPT 5.6 Sol and Claude Fable 5, in overall user experience and some general evaluations. However, the company reports competitive results across coding, terminal operation, software engineering, multimodal tasks, and long duration agent workflows. These benchmark results remain largely based on Moonshot’s own evaluations and should be independently tested as wider access becomes available.
The most technically ambitious Kimi K3 demonstration involved the autonomous design of a specialized accelerator for a smaller model based on the same architecture. During a continuous 48 hour run, Kimi K3 handled the design, optimization, timing closure, verification, and simulation process using open EDA tools and the Nangate 45 nm standard cell library.
The resulting design occupies less than 4 mm², operates at 100 MHz, and contains approximately 1.46 million standard cells, 0.277 MB of SRAM, and an INT4 multiply and accumulate array with integrated dequantization. Moonshot reports simulated decoding performance of more than 8,700 tokens per second, with the published design reaching 8,721 tokens per second.
It is important to distinguish this achievement from manufacturing a finished processor. Moonshot has demonstrated a completed and verified digital chip design operating in simulation. The company has not announced that the design was fabricated, packaged, tested as physical silicon, or prepared for commercial production. The experiment is therefore best understood as evidence of autonomous semiconductor engineering capability rather than a completed consumer or data center product.
The chip project was not Kimi K3’s only major engineering demonstration. Moonshot also asked the model to create a GPU programming system from the ground up. The resulting project, called MiniTriton, includes a domain specific language, an intermediate representation built around MLIR, optimization passes, PTX code generation, and a runtime environment. Moonshot claims MiniTriton matched or exceeded NVIDIA Triton and PyTorch compilation performance in selected supported workloads while successfully running nanoGPT training with stable convergence.
Kimi K3’s highlighted autonomous capabilities include:
Chip design: Created and verified a specialized AI accelerator during a 48 hour autonomous workflow
GPU compiler development: Built MiniTriton with its own compiler pipeline and Tensor Core support
Kernel optimization: Profiled, rewrote, and benchmarked GPU kernels across NVIDIA H200 hardware and another general purpose GPU platform
Scientific research: Reviewed more than 20 research papers, generated more than 3,000 lines of code, evaluated more than 300 equations of state, and produced an interactive astrophysics dashboard
Video production: Edited its own promotional video from 56 source clips with synchronized cuts, audio processing, and multiple revision cycles
Motion graphics: Produced an animated technical explanation of its architecture inspired by the visual presentation style associated with 3Blue1Brown
Interactive content: Converted images, concepts, and videos into early playable experiences through combined visual analysis and code generation
For game development, Moonshot says Kimi K3 can use images and video references to generate interactive 3D environments while continuously examining screenshots and modifying its code. One demonstration presents a character riding a horse through a rural environment containing vegetation, water, buildings, and dynamic rain. The results remain early, but they show how multimodal agents could rapidly create prototypes from visual reference material.
Kimi K3 combines strong 3D reasoning, coding, and vision capabilities to turn concepts, images, and videos into fully playable interactive experiences.
— Kimi.ai (@Kimi_Moonshot) July 16, 2026
Kimi K3 achieves true "vision in the loop" by seamlessly iterating between code and live screenshots pic.twitter.com/E9nw6yivcT
That capability arrives while the gaming industry is already experimenting with AI generated worlds and autonomous coding agents. Google Project Genie, can generate temporary interactive environments but is still positioned as research rather than a complete game creation platform. Another Experiment involving an AI generated GTA style project demonstrated how quickly autonomous agents can produce roads, vehicles, environments, and gameplay systems, while also revealing serious problems with coordination, technical consistency, and creative direction.
Kimi K3 could give creators a more integrated workflow by combining visual understanding, software engineering, terminal operation, 3D reasoning, and continuous iteration inside one model. However, generating a playable prototype remains very different from producing a complete commercial game. Large projects still require intentional game design, reliable performance, original art direction, animation, narrative development, quality assurance, security, licensing compliance, and extensive human supervision.
The model also carries substantial infrastructure requirements. Moonshot recommends deploying Kimi K3 across supernode configurations containing at least 64 accelerators. Its 2.8 trillion parameter architecture, extreme expert count, and 1 million token context window make local deployment unrealistic for normal consumer PCs and most independent studios, even with its MXFP4 weight format. Hosted APIs and shared infrastructure will therefore remain the most practical access routes for many developers.
Kimi K3’s chip experiment is impressive, but the strongest part of the demonstration is not the simulated figure of 8,721 tokens per second. The more significant development is that one model coordinated architecture design, hardware description, optimization, timing, verification, and simulation across a continuous 48 hour engineering session.
This points toward AI becoming an active engineering participant rather than simply a coding assistant. A capable agent could eventually explore thousands of accelerator configurations, identify performance bottlenecks, rewrite supporting software, and produce workload specific hardware proposals before human engineers select the most practical design.
The gaming implications are equally important but should not be exaggerated. Kimi K3 could dramatically accelerate prototyping, level blockouts, interface development, gameplay experiments, technical art workflows, and automated testing. It does not eliminate the need for developers. Instead, it expands how much a small team can explore before committing production resources.
Moonshot’s biggest challenge will be proving these demonstrations outside controlled internal environments. Independent researchers will need access to the full weights, technical report, source materials, compiler project, chip design files, benchmark methodology, and reproducible testing conditions before Kimi K3’s most ambitious claims can be evaluated properly.
Would you trust an AI model to design specialized gaming or computing hardware, or should autonomous chip development always remain under direct human supervision?
