-
Running Multiple Local LLMs: Model Switching and Routing
Run several local LLMs on one GPU: compare llama.cpp router mode, llama-swap, and LiteLLM for on-demand model switching, routing, and cloud fallback.
-
Benchmarking Local LLMs: Tokens per Second Across Hardware
Measure real local LLM speed: tokens per second and time to first token across GPUs, Apple Silicon, and CPU, plus how to benchmark your own setup.
-
Local Voice Assistants: Build a Privacy-First Alexa Alternative
Build a fully local voice assistant with Wyoming, Whisper, Piper, openWakeWord, Home Assistant, and Ollama — no cloud, no data leaving your home network.
-
Apple Neural Engine Explained: Optimizing Models for Apple Silicon
How the Apple Neural Engine works, which operations actually reach it, how to check compute unit assignment in Xcode, and how to design Core ML models it runs.
-
Edge Impulse Tutorial: No-Code TinyML for Embedded Devices
Train and deploy TinyML models without code using Edge Impulse: data collection, impulse design, the EON Tuner, validation, and deploying to Arduino or a browser.
-
MediaPipe Tutorial: Real-Time On-Device ML for Vision and Audio
Run real-time vision and audio ML on-device with MediaPipe: setup for Python, Android, and web, gesture and audio pipelines, Model Maker, and pitfalls.
-
Core ML Tutorial: Deploy Machine Learning Models on iPhone
Deploy machine learning models on iPhone with Core ML: convert PyTorch with coremltools, integrate in Xcode, run Vision predictions, and cut size.
-
Fine-Tuning a Local LLM with LoRA on Consumer Hardware
Fine-tune a local LLM with LoRA and QLoRA on a consumer GPU: hardware sizing, Unsloth setup, dataset prep, training, evaluation, and running it locally.
-
LM Studio Tutorial: Run LLMs Locally with a GUI
Install LM Studio and run LLMs locally: download quantized GGUF or MLX models, chat with documents offline, and serve an OpenAI-compatible API on your machine.
-
Best Books on Deep Reinforcement Learning for Robotics
Four deep RL books compared for robotics work: Sutton & Barto, Lapan, Morales, and Kober & Peters — plus the reading order that actually works.