All posts

local ai

4 posts

Perplexity Releases Lily for Local Qwen 35B Inference on Apple Silicon
Open SourceSep 3, 2026·5 min read

Perplexity Releases Lily for Local Qwen 35B Inference on Apple Silicon

Perplexity has introduced Lily, a Metal inference server optimized for running Qwen 35B models on Apple hardware with advanced session caching.

Read more →
Google Launches Gemma 4 12B: Encoder-Free Multimodal AI for Laptops
AnnouncementsJun 4, 2026·4 min read

Google Launches Gemma 4 12B: Encoder-Free Multimodal AI for Laptops

Google's new Gemma 4 12B brings advanced reasoning and native audio/vision capabilities to 16GB laptops using a novel encoder-free architecture.

Read more →
Memory OS: The 6-Layer Local Memory Stack Curing Hermes Agent's Amnesia
AgentsJun 1, 2026·3 min read

Memory OS: The 6-Layer Local Memory Stack Curing Hermes Agent's Amnesia

Discover how Memory OS, a 6-layer local memory infrastructure by Claudio Drews, gives Hermes Agent permanent memory, structured facts, and semantic search.

Read more →
llama.cpp Adds Multi-Token Prediction for Qwen3.6: A Massive Speed Boost for Local AI
AI ToolsMay 21, 2026·4 min read

llama.cpp Adds Multi-Token Prediction for Qwen3.6: A Massive Speed Boost for Local AI

llama.cpp has merged Multi-Token Prediction (MTP) support for the Qwen3.6 model family, with Georgi Gerganov calling it a 'significant milestone for the local AI ecosystem.' The change enables up to 2.5x faster inference on commodity hardware, with no additional model required — just three extra flags at run time.

Read more →