Learn Machine Learning, MLOps & LLMs with Real Code
Step-by-step machine learning and AI tutorials — MLOps, LLMs, RAG, reinforcement learning and local AI, with working Python code you can run today.
-
Blue-Green Deployments for ML Models
Blue-green deployment for ML models: Kubernetes selector swaps, Argo Rollouts automation, and an instant rollback path when a release goes wrong in production.
-
GitOps for Machine Learning: Version-Controlled Deployments
GitOps for ML: a complete guide to ArgoCD vs Flux for model deployments, DVC data versioning, canary rollouts, and deterministic rollbacks with git revert.
-
Docker Compose for ML Projects: Multi-Container Development Setup
Docker Compose for ML projects: set up GPU-enabled multi-container development with health checks, watch mode, profiles, and resource limits.
-
SageMaker Pipelines Tutorial: End-to-End MLOps on AWS
Build reproducible, auditable ML workflows on AWS: SageMaker Pipelines tutorial covering the @step decorator, step classes, and FailStep gates.
-
Best RAM and Storage Upgrades for Local AI Workstations
Best RAM and storage upgrades for local AI workstations in 2026: how much memory 8B-14B models need, NVMe sizing, and how to shop through the DRAM shortage.
-
Best Laptops for Running Local LLMs and Edge AI Development
Best laptops for local LLMs in 2026: why memory bandwidth beats GPU name, how much memory 7B to 120B models need, and Apple vs NVIDIA vs AMD picks.
-
Speculative Decoding Explained: Speed Up Local LLM Inference
Speculative decoding explained: how draft models, n-gram lookup, Medusa, and EAGLE speed up local LLM inference 2-3x with mathematically identical output quality.
-
Offline AI Apps: Building Machine Learning Apps That Work Without Internet
Build offline AI apps: how WebLLM, Transformers.js, and ONNX Runtime Web run machine learning in the browser with WebGPU acceleration and WASM fallback.
-
Running Multiple Local LLMs: Model Switching and Routing
Run several local LLMs on one GPU: compare llama.cpp router mode, llama-swap, and LiteLLM for on-demand model switching, routing, and cloud fallback.
-
Benchmarking Local LLMs: Tokens per Second Across Hardware
Measure real local LLM speed: tokens per second and time to first token across GPUs, Apple Silicon, and CPU, plus how to benchmark your own setup.
Never miss an article
Follow the blog via RSS — new tutorials as they publish.