Small language models (SLMs) are AI text models with roughly 1 million to 10 billion parameters, according to Hugging Face’s 2025 overview, small enough to run on a phone, a laptop, or one server GPU. They extract fields and classify text at a fraction of a large language model’s serving cost. They also fail in a predictable way: they lose facts before they lose the ability to reason. This guide covers how small language models work, which ones matter in 2026, where they beat larger systems, and the point where you should switch back to an LLM.
What Small Language Models Are and Where “Small” Ends
Small language models are transformer models with about 1 million to 10 billion parameters, per Hugging Face’s 2025 overview, against hundreds of billions in large language models.
A parameter is one adjustable number the model learns during training. More parameters give a model more room to store patterns, and they also cost more memory and energy every time it answers.
IBM’s 2026 explainer describes SLMs as smaller in scale and scope than LLMs. The design itself is the same: SLMs use the transformer architecture found in every generative pre-trained transformer model.
The cutoff moves. Peter Belcak, lead author of NVIDIA’s 2025 agentic AI paper, said in a 2025 Arize interview that his team treats models below roughly 10 billion parameters as small. Some practitioners dislike the label, since a billion parameters is hardly tiny, as Hugging Face’s 2025 overview points out.
Scale shifted fast. Microsoft’s 2024 Phi-3 report points to GPT-2’s 1.5 billion parameters in 2019 as the starting point of the large-model climb. The same report describes a 3.8-billion-parameter model that runs on a phone.
Small Language Models Examples Worth Knowing in 2026
Phi-3-mini has 3.8 billion parameters and scored 68.8 on the MMLU knowledge benchmark in Microsoft’s 2024 report, against 71.4 for GPT-3.5.
The table lists models with published specs. Each row names its source year, since specs in this field go stale within months.
| Model | Maker | Parameters | Reported spec | Source |
|---|---|---|---|---|
| Phi-3-mini | Microsoft | 3.8 billion | 68.8 MMLU; 1.8 GB at 4-bit | Microsoft, 2024 |
| Phi-3-small | Microsoft | 7 billion | 75.3 MMLU | Microsoft, 2024 |
| Phi-3-medium | Microsoft | 14 billion | 78.2 MMLU (preview) | Microsoft, 2024 |
| Apple on-device model | Apple | About 3 billion | 2-bit quantization-aware training | Apple, 2025 |
| SmolLM2 | Hugging Face | 1.7 billion | Trained on FineMath, Stack-Edu, SmolTalk | Hugging Face, 2025 |
| Granite 3.0 | IBM | 2 billion and 8 billion | Base and instruction-tuned versions | IBM, 2026 |
Google DeepMind’s Gemma 3n family targets the same on-device niche, according to BentoML’s 2026 roundup of open-source SLMs.
Look at the jump from 3.8 billion to 7 billion parameters: 6.5 MMLU points in Microsoft’s 2024 tests, for about 1.8 times the memory. Whether that trade pays off depends on how much RAM your device has.
How Small Language Models Fit on a Phone: The Memory Math
Microsoft’s 2024 Phi-3 technical report shows a 3.8-billion-parameter model, quantized to 4 bits, using 1.8 GB of memory and generating over 12 tokens per second on an iPhone 14.
A token is a word fragment, so 12 per second is roughly nine words a second, faster than most people read. Quantization means storing each weight with fewer bits. The Phi-3 team trained in bfloat16, which uses 16 bits, or 2 bytes, per weight. At that precision 3.8 billion weights take about 7.6 GB. At 4 bits they take about 1.9 GB, which matches the reported 1.8 GB.
That gives you a rule of thumb: divide the parameter count in billions by 2 to get gigabytes at 4-bit, then add room for the context cache. A 7-billion-parameter model needs about 3.5 GB for weights. If you run it on a laptop with 8 GB of RAM, roughly 4.5 GB remains for the operating system and your prompt.
Lower precision costs accuracy. According to Apple Machine Learning Research, its roughly 3-billion-parameter on-device model used 2-bit quantization-aware training in 2025, which retrains the model with the compression in place to recover lost quality.
A 14-billion-parameter model needs about 7 GB at 4 bits, which leaves little room on an 8 GB device.
Where Small Language Models Beat Large Ones
NVIDIA researchers estimate in their 2025 paper that serving a 7-billion-parameter small language model costs 10 to 30 times less than serving a 70 to 175 billion parameter LLM.
The paper counts the saving in latency, energy use, and floating-point operations. Its argument targets agentic AI, where a model performs a small number of specialized tasks repeatedly with little variation.
Typical jobs of that kind: pulling an order number from an email, tagging a support ticket, choosing which tool an agent calls next. None of them needs a model that can also discuss philosophy.
Read the number with care. The paper is a position piece rather than a new benchmark, as Arize’s 2025 write-up of the authors’ talk describes it, and the 10 to 30 times figure comes from earlier studies the authors cite. Measure your own workload before you build a budget around it.
If most of your requests share one shape, that price gap is your business case.
Small Language Models Fail on Facts Before They Fail on Reasoning
Phi-3-mini scored 82.5 on the GSM-8K math benchmark, above GPT-3.5’s 78.1, yet only 64.0 on the TriviaQA fact benchmark against GPT-3.5’s 85.8, per Microsoft’s April 2024 report.
The split follows from how Microsoft built the model. Its 2024 report says the team filtered training data to leave capacity for reasoning, and names a Premier League match result as the kind of data it removed.
The report’s weakness section says phi-3-mini lacks the capacity to store much factual knowledge, and it suggests pairing the model with a search engine. Figure 4 in the report shows the same model’s answer with and without search, side by side.
Two practical consequences follow. A small model reading a document you paste into the prompt works from facts you supplied, so its weak recall matters less. A small model asked to recall a product spec, a statute number, or last night’s score has to draw on memory it lacks, and that’s where invented answers come from.
Which of your tasks needs the model to know something, and which only needs it to read something?
When a Small Language Model Is the Wrong Choice
Microsoft’s 2024 report gives Phi-3-mini a default 4K-token context window, so long contracts and codebases push you toward a larger model or its separate 128K variant.
Size inside the small range doesn’t guarantee gains either. Phi-3-medium scored 55.5 on the HumanEval coding benchmark against 59.1 for Phi-3-small in Microsoft’s 2024 numbers, and the report labels the medium figures a preview. Microsoft says some benchmarks improve much less from 7 to 14 billion parameters than from 3.8 to 7 billion.
Language coverage is a third limit. The same report says Microsoft mostly restricted phi-3-mini to English and calls multilingual ability an important next step. If your users write in Hindi or Portuguese, test before you commit.
Open-ended work is the fourth. NVIDIA’s 2025 paper argues for small models on repetitive agent tasks and leaves general conversation to large ones, and it says mixing model sizes suits how agentic tasks vary. The usual pattern is a small model as the default with a large model behind it for the hard calls.
So the answer to “small or large” is often both, and the rule that routes between them matters more than either model.
Running Small Language Models Locally: Phones, Cars, and Factories
Apple’s roughly 3-billion-parameter on-device model, described in its 2025 tech report, runs on Apple silicon itself, so requests it handles skip the server round trip.
Microsoft Azure’s overview says SLMs can work without constant cloud connectivity, and it lists offline translation and virtual assistants as examples. Example: per The Conversation’s August 2026 explainer, Microsoft’s Phi-3 models help power an agricultural information platform in India that serves farmers in remote places with limited internet.
Vehicles face the same constraint. IBM’s 2026 overview lists vehicle navigation assistance as a use, since a compact model can run on a car’s onboard computers. Cars split the work the same way in advanced driver assistance system integration, where edge computing handles real-time tasks like braking and steering while cloud services handle non-critical updates.
Factories share the constraint too. In industrial IoT sensor deployment, edge devices process data locally and send only important results to the cloud, which is where a small language model can sit when a gateway needs to read text or logs.
The question for your deployment is where the weakest connection sits, because the model has to live there.
How to Choose a Small Language Model
Estimate the memory a small language model needs by dividing its parameters in billions by 2 for 4-bit weights, so a 7-billion-parameter model needs about 3.5 GB before context and overhead.
Then work through six checks in order:
- Sort tasks. Mark each task as “needs recall” or “needs reading.” Reading tasks suit small models.
- Check memory. Apply the divide-by-2 rule against the RAM on your target device.
- Check context. Compare your longest input with the model’s window. Phi-3-mini’s default is 4K tokens (Microsoft, 2024).
- Test 50 real examples. Score a small model and an LLM on the same 50 inputs from your own data, then compare the gap with your error tolerance.
- Add retrieval. For recall tasks, supply facts through search, the fix Microsoft’s 2024 report suggests.
- Keep a fallback. Send requests the small model flags as uncertain, or that fail a format check, to a larger model.
For practitioners with two or more years in production, NVIDIA’s 2025 paper describes a conversion path that starts with logging agent calls and clustering them by task. Within 30 days you can log two weeks of calls, cluster them, and fine-tune one small model on the largest cluster while the large model stays as the fallback.
People Also Ask
What is the difference between a small language model and a large language model?
A small language model has roughly 1 million to 10 billion parameters, while a large language model has hundreds of billions, per Hugging Face’s 2025 overview. Both use transformers. Small models run on phones and laptops and suit narrow tasks. Large models need data-center hardware and handle open-ended work, broad knowledge, and long documents better.
What are examples of small language models?
Examples include Microsoft’s Phi-3-mini (3.8 billion parameters), Apple’s on-device model (about 3 billion), Hugging Face’s SmolLM2 (1.7 billion), and IBM’s Granite 3.0 (2 and 8 billion), per each maker’s 2024 to 2026 publications. Google DeepMind’s Gemma 3n targets phones too. Check each vendor’s model card for current sizes, because versions change quickly.
Can small language models run on a phone?
Yes. Microsoft’s 2024 Phi-3 report ran a 4-bit, 3.8-billion-parameter model on an iPhone 14 at over 12 tokens per second, using 1.8 GB of memory. Apple’s 2025 report describes a roughly 3-billion-parameter model built for Apple silicon. Bigger models need more RAM: about 7 GB for 14 billion parameters at 4 bits.
Are small language models better than LLMs?
Small language models beat LLMs on cost and latency for narrow, repetitive tasks, and lose on broad knowledge. NVIDIA’s 2025 paper estimates serving costs 10 to 30 times lower at 7 billion parameters. Microsoft’s 2024 report puts phi-3-mini ahead of GPT-3.5 on GSM-8K math (82.5 versus 78.1) and behind on TriviaQA (64.0 versus 85.8).
What are small language models used for?
Small language models handle classification, field extraction, on-device assistants, offline translation, and repeated agent steps such as routing requests to tools. IBM’s 2026 overview also lists sentiment analysis and vehicle navigation assistance. They work best when the task supplies its own facts, since Microsoft’s 2024 report says they store little factual knowledge.
Frequently Asked Questions
How much cheaper is a small language model to run?
NVIDIA’s 2025 paper estimates that serving a 7-billion-parameter model costs 10 to 30 times less than serving a 70 to 175 billion parameter model, measured in latency, energy, and floating-point operations. That’s an estimate drawn from earlier studies, so your ratio will vary with hardware, batch size, and prompt length. Self-hosting also adds fixed costs: a GPU or device to run it on, and someone to maintain it. Run a two-week pilot on real traffic and compare the cost per thousand requests before you commit.
How do you fine-tune a small language model?
Most teams use parameter-efficient fine-tuning, which trains a small set of added weights while the base model stays frozen. LoRA, introduced by Hu et al. in 2021, does this directly, and QLoRA, introduced by Dettmers et al. in 2023, does the same on a 4-bit base model. A 2024 arXiv study on agent models used both. Start with a small, clean set of examples for one task, hold out 50 for testing, and compare against the untuned model. If tuning doesn’t beat a good prompt, skip it.
Do small language models hallucinate?
Yes. Microsoft’s 2024 report says factual inaccuracies remain a challenge for phi-3-mini. In Microsoft’s in-house test, phi-3-mini-4k scored 0.603 for ungroundedness on a 0 to 4 scale where lower is better, against 1.481 for Phi-2 and 0.935 for Mistral 7B. Newer small models improve, but the risk stays highest on recall questions. Supply source text in the prompt, ask for answers drawn only from it, and spot-check the outputs.
Are small language models safe for private data?
Running a model locally keeps prompts off third-party servers, which removes one exposure point. It doesn’t remove model-level risk. Microsoft’s 2024 in-house test found a 12.29% jailbreak defect rate for phi-3-mini-4k, meaning about 12 in 100 adversarial samples produced a response of at least the lowest harm severity. Treat a local model like any other component: filter inputs, restrict tool access, and log outputs.
What hardware do you need to run a small language model locally?
Memory decides most of it. Divide the parameter count in billions by 2 to get gigabytes at 4-bit, then add room for the context cache. Microsoft’s 2024 report measured 1.8 GB for a 3.8-billion-parameter model on an iPhone 14. A 7-billion-parameter model needs about 3.5 GB for weights, so a laptop with 8 GB of RAM can probably hold it. Speed then depends on the chip, and Microsoft reported over 12 tokens per second on the A16 Bionic.
KEEP READING
AI ethics and bias mitigation mean defining fair outcomes, testing for gaps, and fixing them across the three stages in NIST's 2022 SP 1270. Those stages are pre-design, design and [...]
An AI governance framework is the set of policies, risk tiers, and approval gates that decides which AI systems your organization can ship and which ones need a human sign-off [...]
AI agents and chatbots get sold as the same thing, and that's costing companies money before they've deployed anything. A chatbot matches what you type to a pre-written answer and [...]
Generative AI tools for content creation now range from a general chatbot that drafts an email in ten seconds to a specialized platform that holds a brand's tone across a [...]
AI personalized learning in K-12 EdTech means software that changes what a student sees next based on how they just answered, not a fixed chapter everyone works through at the [...]