
Perplexity Releases Lily for Local Qwen 35B Inference on Apple Silicon
Perplexity has introduced Lily, a Metal inference server optimized for running Qwen 35B models on Apple hardware with advanced session caching.
Read more →
Google Launches Gemma 4 12B: Encoder-Free Multimodal AI for Laptops
Google's new Gemma 4 12B brings advanced reasoning and native audio/vision capabilities to 16GB laptops using a novel encoder-free architecture.
Read more →
Memory OS: The 6-Layer Local Memory Stack Curing Hermes Agent's Amnesia
Discover how Memory OS, a 6-layer local memory infrastructure by Claudio Drews, gives Hermes Agent permanent memory, structured facts, and semantic search.
Read more →
llama.cpp Adds Multi-Token Prediction for Qwen3.6: A Massive Speed Boost for Local AI
llama.cpp has merged Multi-Token Prediction (MTP) support for the Qwen3.6 model family, with Georgi Gerganov calling it a 'significant milestone for the local AI ecosystem.' The change enables up to 2.5x faster inference on commodity hardware, with no additional model required — just three extra flags at run time.
Read more →