Back to News
CohereSeptember 9, 2026

Megakernel Serving Engine

Open Source
Inference Engine
CUDA
LLM Serving
H100

Explore Megakernel Serving Engine

Visit the official website to learn more and get started

### TL;DR

The Megakernel Serving Engine is a high-performance inference system designed specifically for the North Mini Code model. It utilizes a persistent CUDA megakernel to execute the complete decode forward pass, significantly reducing overhead by eliminating per-op kernel launches and grid synchronization.

Key Insights & Metrics

Pricing
Open Source (Apache-2.0 license)
Cost structure
Version
0.1.0
Current release version
Hardware
Single NVIDIA H100 (sm_90a), CUDA 13+, Linux
Compute requirements
Category
Open Source
Licensing model
Region
Canada
Primary region

Key Features

  • Persistent CUDA megakernel for complete decode forward pass
  • OpenAI-compatible API for completions and chat-completions
  • Continuous batching and paged KV cache support
  • Aggressive SM backfilling for MoE tiles
  • Support for sliding-window attention and prefix caching
Ad
Master AI Marketing: Work Smarter, Not Harder

Master AI Marketing: Work Smarter, Not Harder

Unlock the power of AI to automate workflows and scale your results. Join 1.5M+ professionals and stay ahead of the curve with the latest AI insights.

Sponsored

Discussion

0
Upvotes
0
Downvotes
0 reviews

Sign in to leave a review

Reviews

No reviews yet. Be the first to review!