Back to News
Cohere•September 9, 2026
Megakernel Serving Engine
Open Source
Inference Engine
CUDA
LLM Serving
H100
### TL;DR
The Megakernel Serving Engine is a high-performance inference system designed specifically for the North Mini Code model. It utilizes a persistent CUDA megakernel to execute the complete decode forward pass, significantly reducing overhead by eliminating per-op kernel launches and grid synchronization.
Key Insights & Metrics
Pricing
Open Source (Apache-2.0 license)
Cost structure
Version
0.1.0
Current release version
Hardware
Single NVIDIA H100 (sm_90a), CUDA 13+, Linux
Compute requirements
Category
Open Source
Licensing model
Region
Canada
Primary region
Key Features
- Persistent CUDA megakernel for complete decode forward pass
- OpenAI-compatible API for completions and chat-completions
- Continuous batching and paged KV cache support
- Aggressive SM backfilling for MoE tiles
- Support for sliding-window attention and prefix caching
Ad

Master AI Marketing: Work Smarter, Not Harder
Unlock the power of AI to automate workflows and scale your results. Join 1.5M+ professionals and stay ahead of the curve with the latest AI insights.
Sponsored
Discussion
0
Upvotes
0
Downvotes
0 reviews
Sign in to leave a review
Reviews
No reviews yet. Be the first to review!