{
 "sym": "Sm",
 "updated": "2026-08-06",
 "verdict": "Ollama, if you can only hold one \u2014 8.9M monthly developers, presence in 85% of the Fortune 500, and a $65M Series B (Jul 2026) make it the de-facto local runtime, and local inference stays free and unlimited. Gemma 4 E-series is the model to load first: Apache-2.0, multimodal, and built for 2\u20134GB footprints. LM Studio wins when you want a GUI, Mac-optimized MLX speed, and free commercial use; llama.cpp when inference ships inside your own product; Qwen3.5 Small when your edge agent needs eyes and ears at under 1GB.",
 "top5": [
  {
   "rank": 1,
   "name": "Ollama",
   "maker": "Ollama Inc.",
   "url": "https://ollama.com",
   "docs": "https://docs.ollama.com",
   "pricing": "Local: free, unlimited \u00b7 Cloud free tier \u00b7 Pro $20/mo \u00b7 Max $100/mo (signups paused)",
   "best_for": "The default way to pull and run any open model locally \u2014 one command, OpenAI-compatible API, and every agent framework already speaks it.",
   "why": "The category's Docker moment: 8.9M monthly developers, ~176k GitHub stars, used inside 85% of the Fortune 500, all built by a 14-person team that raised a $65M Series B in July 2026 ($88M total). In 2026 it re-platformed Apple Silicon inference onto MLX \u2014 up to 90% faster for coding agents (Jun 2026) \u2014 while keeping GGUF/llama.cpp compatibility.",
   "watch": "The cloud pivot ($20\u2013100/mo tiers, datacenter models) is where monetization pressure lives \u2014 local stays free today, but the incentive gradient now points at hosted usage. Power users note the new app and MLX default reduced the old single-binary simplicity.",
   "evidence": [
    {
     "stat": "$65M Series B led by Theory Ventures; 8.9M monthly developers; 85% of Fortune 500 (Jul 9, 2026)",
     "src": "https://techcrunch.com/2026/07/09/popular-open-source-ai-developer-tool-ollama-raises-65m-grows-to-nearly-9m-users/"
    },
    {
     "stat": "MLX engine on Apple Silicon: up to 90% faster for coding agents with Gemma 4 (Jun 29, 2026)",
     "src": "https://ollama.com/blog"
    },
    {
     "stat": "Local models always unlimited and free; cloud Pro $20/mo, Max $100/mo (verified Aug 2026)",
     "src": "https://ollama.com/cloud"
    }
   ],
   "tile_note": "local runtime of choice"
  },
  {
   "rank": 2,
   "name": "Gemma 4 (E2B / E4B)",
   "maker": "Google",
   "url": "https://ai.google.dev/gemma",
   "docs": "https://ai.google.dev/gemma/docs",
   "pricing": "Free \u2014 Apache 2.0 weights (license upgraded from the restrictive Gemma terms)",
   "best_for": "The first model to load on any laptop or phone \u2014 multimodal (text, image, audio), 128K context, 2\u20134GB effective footprint.",
   "why": "The on-device workhorse of 2026: released Apr 2, 2026 under Apache 2.0 \u2014 dropping the old Gemma license restrictions \u2014 with E2B (2.3B effective) and E4B (4.5B) variants that run fully offline on mobile and IoT hardware. The family has passed 400M downloads with 100k+ community variants, and E2B hits 60% MMLU-Pro \u2014 numbers that needed 30B+ models two years ago.",
   "watch": "Google's release cadence makes any Gemma pick obsolete in ~12 months, and the 26B/31B variants pull you out of small-model territory (see Ow). Audio input is limited to the smaller variants; long-context quality degrades past 32K on E2B in community testing.",
   "evidence": [
    {
     "stat": "Gemma 4 released Apr 2, 2026, Apache 2.0, five sizes 2.3B\u201331B; family has 400M+ downloads, 100k+ variants",
     "src": "https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/"
    },
    {
     "stat": "E2B: 60% MMLU-Pro at 2.3B effective params; 128K context; text+image+audio on-device (Apr 2026)",
     "src": "https://huggingface.co/blog/gemma4"
    },
    {
     "stat": "Day-0 support across llama.cpp, MLX, WebGPU; runs completely offline on mobile and IoT",
     "src": "https://huggingface.co/blog/gemma4"
    }
   ],
   "tile_note": "the on-device workhorse"
  },
  {
   "rank": 3,
   "name": "LM Studio",
   "maker": "Element Labs",
   "url": "https://lmstudio.ai",
   "docs": "https://lmstudio.ai/docs",
   "pricing": "Free incl. commercial use \u00b7 optional cloud inference $0.13\u20133.00/M tokens \u00b7 Enterprise custom",
   "best_for": "Teams that want local models with a real GUI \u2014 model discovery, chat, RAG, plus Python/TypeScript SDKs and a headless server mode underneath.",
   "why": "The most polished way to run GGUF and MLX models on a desktop, free even for commercial use. 2026 closed its two gaps with Ollama: llmster headless server mode (Jan 2026) for deployments, and the Bionic agent app (Jul 16, 2026) that puts an agentic layer over the same local runtime. Dual-engine (llama.cpp + MLX) with KV-cache checkpointing tuned for long-context agent loops.",
   "watch": "The core app is closed-source \u2014 you are trusting a VC-backed company's roadmap, and the Bionic launch signals focus shifting toward a cloud-billed agent product ('Bionic Pass' pricing still unannounced Aug 2026). CLI repo has ~5.1k stars vs Ollama's 176k \u2014 a fraction of the ecosystem gravity.",
   "evidence": [
    {
     "stat": "Free tier $0 incl. local LLMs and voice transcription; cloud pay-as-you-go $0.13\u20133.00/M tokens (verified Aug 2026)",
     "src": "https://lmstudio.ai/pricing"
    },
    {
     "stat": "Bionic agent app launched Jul 16, 2026 on the LM Studio local runtime; classic app continues",
     "src": "https://lmstudio.ai/blog/introducing-lm-studio-bionic"
    },
    {
     "stat": "llmster headless mode (Jan 2026) + MLX/llama.cpp dual runtime; v0.4.19 by Jul 2026",
     "src": "https://www.kunalganglani.com/blog/lm-studio-vs-ollama"
    }
   ],
   "tile_note": "local llms, real gui"
  },
  {
   "rank": 4,
   "name": "Qwen3.5 Small",
   "maker": "Alibaba",
   "url": "https://qwen.ai",
   "docs": "https://qwen.readthedocs.io",
   "pricing": "Free weights on Hugging Face / ModelScope (Instruct + Base)",
   "best_for": "Multimodal edge agents \u2014 the 0.8B and 2B run on phones and IoT chips with native image/video input; the 4B is the strongest small agent base of 2026.",
   "why": "Released Mar 2, 2026 as a purpose-built small family (0.8B / 2B / 4B / 9B) on the same native-multimodal Qwen3.5 foundation \u2014 a 0.8B model that processes video is the clearest marker yet of the edge-AI era. The 4B adds 262K context and on-demand thinking across 201 languages; the 9B closes on models 5\u201310x its size via scaled RL.",
   "watch": "License terms for the Small series weren't confirmed on a vendor page at review time (the Qwen3 line was Apache 2.0 \u2014 verify per model card). China provenance still blocks adoption in some Western enterprises regardless of open weights, and benchmark claims are mostly self-reported.",
   "evidence": [
    {
     "stat": "Qwen3.5 Small series (0.8B\u20139B) released Mar 2\u20133, 2026; native multimodal from 4B up, edge-tuned below",
     "src": "https://www.marktechpost.com/2026/03/02/alibaba-just-released-qwen-3-5-small-models-a-family-of-0-8b-to-9b-parameters-built-for-on-device-applications/"
    },
    {
     "stat": "Qwen3.5-4B: image+video input, on-demand thinking, 201 languages, 262K context (2026)",
     "src": "https://tinyweights.dev/posts/best-small-language-models-2026/"
    },
    {
     "stat": "Official line: '0.8B/2B \u2192 tiny, fast, great for edge; 4B \u2192 surprisingly strong multimodal base for lightweight agents' (Mar 2026)",
     "src": "https://x.com/Alibaba_Qwen/status/2028460046510965160"
    }
   ],
   "tile_note": "tiny multimodal agents"
  },
  {
   "rank": 5,
   "name": "llama.cpp",
   "maker": "ggml-org (Georgi Gerganov)",
   "url": "https://github.com/ggml-org/llama.cpp",
   "docs": "https://github.com/ggml-org/llama.cpp/tree/master/docs",
   "pricing": "Free, MIT \u2014 vendor-neutral C/C++, no strings",
   "best_for": "Shipping inference inside your own product \u2014 the MIT-licensed engine that runs GGUF models on effectively every chip made.",
   "why": "The engine underneath the whole category: ~122.7k stars, 21k forks, 10k+ commits, and the GGUF quantization format it defined (1.5-bit to 8-bit) is the lingua franca of local AI. CUDA, Metal, Vulkan, HIP, x86, Apple Silicon, even RISC-V \u2014 plus a built-in OpenAI-compatible server and VLM support, with zero runtime dependencies to license or trust.",
   "watch": "It's a toolkit, not a product \u2014 model management, updates, and UX are on you (that's why Ollama and LM Studio exist). On Apple Silicon, MLX now decodes 1.4\u20131.8x faster, which pushed even Ollama to route safetensors to MLX in 2026.",
   "evidence": [
    {
     "stat": "~122.7k stars, ~21.3k forks, 10,273 commits; active daily development (Aug 2026)",
     "src": "https://github.com/ggml-org/llama.cpp"
    },
    {
     "stat": "SOTA inference 'on a wide range of hardware' \u2014 Apple Silicon, x86, RISC-V; CUDA/HIP/Metal/Vulkan backends; 1.5\u20138-bit quantization",
     "src": "https://github.com/ggml-org/llama.cpp"
    },
    {
     "stat": "MLX decodes 1.4\u20131.8x faster than llama.cpp on Apple Silicon, but llama.cpp keeps the prefill/TTFT edge (Mar 2026 benchmarks)",
     "src": "https://yage.ai/share/mlx-apple-silicon-en-20260331.html"
    }
   ],
   "tile_note": "the engine underneath everything"
  }
 ],
 "matrix": {
  "cols": [
   "Type",
   "License",
   "Runs on",
   "Footprint floor",
   "Multimodal",
   "API surface",
   "NPU / accel"
  ],
  "rows": [
   [
    "Ollama",
    "Runtime + app",
    "MIT (server)",
    "macOS \u00b7 Win \u00b7 Linux \u00b7 Docker",
    "~8GB RAM useful",
    "Via models",
    "OpenAI-compat + native",
    "MLX (Apple) \u00b7 CUDA \u00b7 ROCm"
   ],
   [
    "Gemma 4 E-series",
    "Model family",
    "Apache 2.0",
    "Anywhere GGUF/MLX/LiteRT runs",
    "~2GB (E2B)",
    "Text \u00b7 image \u00b7 audio",
    "Weights only",
    "LiteRT-LM, NPU-ready"
   ],
   [
    "LM Studio",
    "Runtime + GUI + SDKs",
    "Closed core, free use",
    "macOS \u00b7 Win \u00b7 Linux",
    "~8GB RAM useful",
    "Via models",
    "OpenAI-compat \u00b7 py/ts SDK \u00b7 CLI",
    "MLX + llama.cpp engines"
   ],
   [
    "Qwen3.5 Small",
    "Model family 0.8\u20139B",
    "Open weights (verify card)",
    "HF/ModelScope \u2192 GGUF/MLX",
    "<1GB (0.8B)",
    "Native, incl. video",
    "Weights only",
    "Edge-tuned variants"
   ],
   [
    "llama.cpp",
    "C/C++ engine",
    "MIT",
    "Everything incl. RISC-V",
    "<1GB possible",
    "VLM support",
    "OpenAI-compat server \u00b7 CLI",
    "CUDA \u00b7 Metal \u00b7 Vulkan \u00b7 HIP"
   ],
   [
    "Phi-4 family",
    "Model family 3.8\u201315B",
    "MIT",
    "GGUF/ONNX everywhere",
    "~3GB (mini)",
    "Vision variants",
    "Weights only",
    "ONNX / DirectML path"
   ],
   [
    "Foundry Local",
    "OS-level runtime",
    "Free, closed",
    "Win \u00b7 macOS (AS) \u00b7 Linux x64",
    "Device-dependent",
    "Whisper + VLMs in catalog",
    "OpenAI-compat SDKs (4 langs)",
    "WinML: NPU/GPU/CPU auto"
   ],
   [
    "Apple Foundation Models",
    "OS framework + ~3B model",
    "Free, closed",
    "iOS / macOS (26+)",
    "0 \u2014 built into OS",
    "Text-focused",
    "Swift API; any provider since WWDC26",
    "Neural Engine"
   ]
  ]
 },
 "rules": [
  {
   "if": "You just want open models running on your machine today",
   "then": "Ollama \u2014 one command, OpenAI-compatible endpoint, and every framework integrates with it. Ignore the cloud tiers until you need datacenter-size models."
  },
  {
   "if": "You're choosing the model, not the runtime",
   "then": "Gemma 4 E4B as default (Apache 2.0, multimodal, 128K); Qwen3.5-4B when the agent needs image/video input; Phi-4-mini when you want MIT license and math/reasoning at 3.8B."
  },
  {
   "if": "Inference ships inside your product",
   "then": "llama.cpp (MIT, every platform) \u2014 or MLX if you're Apple-only, where it decodes 1.4\u20131.8x faster and M5 neural accelerators cut time-to-first-token 4x."
  },
  {
   "if": "You're building a mobile app",
   "then": "Use the platform stack, not a runtime port: Apple Foundation Models (free ~3B on-device, open to third-party models since WWDC 2026) or Google AI Edge/LiteRT with Gemma 4 E2B; Nexa SDK for Qualcomm NPUs."
  },
  {
   "if": "Your frontier-model bill has line items for classification, extraction, or summarization",
   "then": "Move them to a small local model \u2014 NVIDIA's position paper argues most agentic subtasks are 'repetitive, scoped, non-conversational' and SLM-sized; this is the inference-bill signal the element's tagline is about."
  }
 ],
 "field": [
  {
   "name": "Phi-4 family",
   "maker": "Microsoft",
   "note": "MIT-licensed; Phi-4-mini (3.8B) still tops sub-4B reasoning rankings, Phi-4-reasoning-vision-15B added Mar 2026 \u2014 but no Phi-5, and Gemma/Qwen out-shipped it in 2026 (Phi-5 specs circulating are pre-release speculation)",
   "url": "https://azure.microsoft.com/en-us/products/phi",
   "oss": true,
   "entry": "free (MIT)",
   "status": "active"
  },
  {
   "name": "MLX",
   "maker": "Apple",
   "note": "Apple's array framework, now the engine under both Ollama and LM Studio on Apple Silicon; M5 neural accelerators are designed for its compute patterns (4.06x faster TTFT vs M4, Jan 2026)",
   "url": "https://github.com/ml-explore/mlx",
   "oss": true,
   "entry": "free (MIT)",
   "status": "active"
  },
  {
   "name": "gpt-oss-20b",
   "maker": "OpenAI",
   "note": "Apache 2.0, o3-mini-class in 16GB RAM (Aug 2025) \u2014 the small end of OpenAI's open weights; the 120B sibling belongs to Ow",
   "url": "https://openai.com/index/introducing-gpt-oss/",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "SmolLM3-3B",
   "maker": "Hugging Face",
   "note": "The fully-open small model \u2014 Apache 2.0 with reproducible training data and recipes, 128K context; the transparency benchmark for the class",
   "url": "https://huggingface.co/blog/smollm3",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "LFM2 / Liquid Nanos",
   "maker": "Liquid AI",
   "note": "Non-transformer edge models 230M\u20132.6B; LFM2.5-230M claims wins over models 4x its size at extraction; strong in embedded/automotive deals",
   "url": "https://www.liquid.ai/models",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "Foundry Local",
   "maker": "Microsoft",
   "note": "GA Jun 2, 2026 at Build \u2014 free on-device runtime as native SDK (Python/JS/C#/Rust), WinML auto NPU/GPU/CPU, ~24-model catalog; Windows' answer to Ollama",
   "url": "https://learn.microsoft.com/en-us/azure/ai-foundry/foundry-local/",
   "oss": false,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Apple Foundation Models",
   "maker": "Apple",
   "note": "~3B on-device model free to every iOS/macOS app; WWDC 2026 opened the framework to any LLM provider (incl. MLX models from HF) \u2014 the OS is becoming the runtime",
   "url": "https://developer.apple.com/documentation/foundationmodels",
   "oss": false,
   "entry": "free (OS-bundled)",
   "status": "active"
  },
  {
   "name": "Google AI Edge / LiteRT-LM",
   "maker": "Google",
   "note": "On-device stack for Android/cross-platform + AI Edge Gallery app for running Gemma 4 on phones",
   "url": "https://ai.google.dev/edge",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Nexa SDK",
   "maker": "Qualcomm (acq. Nexa AI)",
   "note": "Day-0 NPU/GPU/CPU runtime for PC, mobile, IoT; Nexa AI acquired by Qualcomm Mar 21, 2026 \u2014 chipmakers buying the local-inference layer",
   "url": "https://github.com/qualcomm/nexa-sdk",
   "oss": true,
   "entry": "free",
   "status": "acquired"
  },
  {
   "name": "Docker Model Runner",
   "maker": "Docker",
   "note": "Models as OCI artifacts, `docker model run`, GA 2025 \u2014 the path of least resistance for teams already in Docker Desktop",
   "url": "https://www.docker.com/products/model-runner/",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Jan",
   "maker": "Menlo Research",
   "note": "Open-source local desktop app (1M+ downloads) with its own Jan-Nano agentic small models; the open-source alternative to LM Studio",
   "url": "https://jan.ai",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "llamafile",
   "maker": "Mozilla.ai",
   "note": "Single-file portable LLM executable; dormant through 2025, revived with v0.10 (Mar 2026) \u2014 rebuilt core, GPU support",
   "url": "https://www.mozilla.ai/open-tools/llamafile",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "LocalAI",
   "maker": "mudler (community)",
   "note": "Self-hosted OpenAI-compatible stack \u2014 LLM, image, audio \u2014 for homelab and on-prem deployments",
   "url": "https://localai.io",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "GPT4All",
   "maker": "Nomic AI",
   "note": "The 2023 pioneer of desktop local AI; releases stopped and GitHub issues openly ask 'is GPT4All dead?' \u2014 Nomic's focus moved to embeddings",
   "url": "https://github.com/nomic-ai/gpt4all",
   "oss": true,
   "entry": "free",
   "status": "fading"
  },
  {
   "name": "ExecuTorch",
   "maker": "PyTorch / Meta",
   "note": "PyTorch's edge/mobile inference runtime for embedding models in apps",
   "url": "https://github.com/pytorch/executorch",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Ministral 3",
   "maker": "Mistral AI",
   "note": "Mistral's small edge line (successor to Ministral 3B/8B), day-0 supported in Nexa SDK \u2014 details thin on vendor pages, verify per card",
   "url": "https://mistral.ai/models",
   "oss": true,
   "entry": "free weights",
   "status": "active"
  },
  {
   "name": "Open WebUI",
   "maker": "Open WebUI Inc.",
   "note": "The dominant self-hosted chat UI over Ollama/OpenAI-compat backends \u2014 UI layer, not a runtime; license moved away from pure BSD in 2025",
   "url": "https://openwebui.com",
   "oss": true,
   "entry": "free",
   "status": "active"
  },
  {
   "name": "Ollama Cloud Max tier",
   "maker": "Ollama Inc.",
   "note": "The $100/mo Max tier paused new signups mid-2026 \u2014 capacity constraints reach even the local-first vendors' cloud arms",
   "url": "https://ollama.com/cloud",
   "oss": false,
   "entry": "paused",
   "status": "sunsetting"
  }
 ],
 "signals": [
  {
   "fact": "SLM market: $0.93B (2025) \u2192 $5.45B (2032), 28.7% CAGR; sub-2B-parameter models the fastest-growing segment",
   "src": "https://www.marketsandmarkets.com/PressReleases/small-language-model.asp"
  },
  {
   "fact": "Ollama: $65M Series B (Jul 9, 2026, Theory Ventures; $88M total), 8.9M monthly developers, 14 employees, in 85% of the Fortune 500",
   "src": "https://techcrunch.com/2026/07/09/popular-open-source-ai-developer-tool-ollama-raises-65m-grows-to-nearly-9m-users/"
  },
  {
   "fact": "Platform consolidation, 2026: Microsoft Foundry Local GA (Jun 2), Apple opens Foundation Models to any provider (WWDC, Jun), Qualcomm acquires Nexa AI (Mar 21) \u2014 every OS and chip vendor now ships a free local runtime",
   "src": "https://byteiota.com/microsoft-foundry-local-ga/"
  },
  {
   "fact": "Gemma family passed 400M downloads with 100k+ variants; Gemma 4 (Apr 2, 2026) moved the family to Apache 2.0",
   "src": "https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/"
  },
  {
   "fact": "NVIDIA research position: small language models are 'the future of agentic AI' \u2014 most agent subtasks don't need frontier calls (arXiv 2506.02153, Jun 2025)",
   "src": "https://arxiv.org/abs/2506.02153"
  },
  {
   "fact": "Apple Silicon became a first-class inference target: Ollama and LM Studio both moved to MLX engines in 2026; M5 neural accelerators cut time-to-first-token 4.06x vs M4 (Apple research, Jan 2026)",
   "src": "https://yage.ai/share/mlx-apple-silicon-en-20260331.html"
  }
 ],
 "notes": "Ranking criteria: developer adoption, model quality per GB, license freedom, and whether the thing still ships \u2014 runtimes and model families judged together because that's how the element is used. Scope: big open-weight families (Llama, DeepSeek, GLM, Kimi, Mistral Large, gpt-oss-120b) \u2192 element Ow; serving infra (vLLM, SGLang) \u2192 Ow/infra; chat UIs (Open WebUI) get a field line only; agent harnesses that happen to run local models \u2192 Ca/Ag. Conflicts resolved: llama.cpp star counts vary by aggregator (103.8k on SEO sites) \u2014 we cite GitHub directly (~122.7k, Aug 2026). Ollama's MLX switch has two dates in the wild \u2014 v0.19 preview Mar 30, 2026 and engine updates Jun 2026; both are real, preview then rollout. 'Phi-5' guides circulating in mid-2026 are explicitly based on pre-release speculation (Spheron admits it); latest real Phi is Phi-4-reasoning-vision (Mar 2026). Qwen3.5 Small license unconfirmed on a vendor page at review time \u2014 flagged in the entry, single-source benchmark claims likewise. LM Studio 'Bionic Pass' pricing unannounced; its user counts are not public (private company, no primary source) so we deliberately cite no user figure. Ollama Cloud Max 'paused' status is from the vendor pricing page and may change weekly.",
 "sources": [
  "https://techcrunch.com/2026/07/09/popular-open-source-ai-developer-tool-ollama-raises-65m-grows-to-nearly-9m-users/",
  "https://ollama.com/cloud",
  "https://ollama.com/blog",
  "https://docs.ollama.com",
  "https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/",
  "https://huggingface.co/blog/gemma4",
  "https://ai.google.dev/gemma/docs",
  "https://www.marktechpost.com/2026/03/02/alibaba-just-released-qwen-3-5-small-models-a-family-of-0-8b-to-9b-parameters-built-for-on-device-applications/",
  "https://lmstudio.ai/pricing",
  "https://lmstudio.ai/blog/introducing-lm-studio-bionic",
  "https://lmstudio.ai/docs",
  "https://github.com/ggml-org/llama.cpp",
  "https://yage.ai/share/mlx-apple-silicon-en-20260331.html",
  "https://arxiv.org/abs/2506.02153",
  "https://www.marketsandmarkets.com/PressReleases/small-language-model.asp",
  "https://byteiota.com/microsoft-foundry-local-ga/",
  "https://dev.to/arshtechpro/wwdc-2026-apple-just-opened-the-foundation-models-framework-to-any-llm-provider-5ejn",
  "https://github.com/qualcomm/nexa-sdk/discussions/1058",
  "https://openai.com/index/introducing-gpt-oss/",
  "https://blog.mozilla.ai/llamafile-reloaded-whats-new-in-v0-10-0/",
  "https://tinyweights.dev/posts/best-small-language-models-2026/",
  "https://en.wikipedia.org/wiki/Phi_(language_model)",
  "https://www.kunalganglani.com/blog/lm-studio-vs-ollama"
 ],
 "element": {
  "number": 3,
  "name": "Small / Local Model",
  "group": "Intelligence",
  "essential": false,
  "edition": "v2026.Q3",
  "revision": "r7",
  "license": "CC BY 4.0 \u2014 cite elems.ai",
  "url": "https://elems.ai/e/sm.html"
 }
}