For the last few years the default advice in AI development has been simple: use the biggest model you can afford. That advice is starting to age. A growing number of business tasks do not need a model that can write poetry, solve math olympiad problems and summarise legal briefs all at once. They need a model that does one job quickly, cheaply and predictably. That is where small language models (SLMs) earn their place.
At Mavani Solution we meet this question in almost every AI discovery call. A founder has a prototype built on a large hosted model, it works, and then the invoice and the response times arrive. The conversation moves from "can AI do this?" to "can AI do this at a price and speed our customers accept?" This guide explains what SLMs are, where they win, where they lose, and how to adopt one without taking on unnecessary risk.
There is no official cut-off, but in practice an SLM is a model small enough to run on a single modest GPU, a CPU server, a laptop or even a phone after quantization. Large models are trained to be generalists. Small models are usually either distilled from a larger teacher model or trained on carefully curated data, and they are often specialised further with fine-tuning for a narrow domain.
The trade-off is straightforward. You give up some breadth and some depth of reasoning. In return you get lower latency, lower cost per request, easier self-hosting and tighter control over where data goes. For a customer support classifier or an invoice field extractor, that trade is often a bargain.
SLMs shine when the task is narrow, the inputs are predictable and the output format is constrained. Typical examples include:
If your use case looks like this, a small model deserves a trial. Our write-up on on-device AI for mobile apps covers the mobile side of the same idea in more depth.
Be honest about the limits. Large models remain the better choice for open-ended reasoning, long multi-step planning, messy multi-document analysis and tasks where the input is unpredictable. If a user can ask anything and you must handle it gracefully, a frontier model is usually safer. Many production systems end up hybrid: an SLM handles the common, easy requests and a larger model handles the escalations.
Treat model size as a dial, not a religion. The right model is the smallest one that clears your quality bar on your own data.
Numbers vary by provider and change often, so treat any figure as a planning aid, not a quote. For example, imagine an app that classifies 500,000 support messages a month. If a hosted frontier model costs several times more per call than a self-hosted small model, the monthly gap could be large enough to fund an engineer's time to build the pipeline. Latency tells a similar story: a small model on nearby hardware could often respond faster than a round trip to a large hosted model, which may matter for features that must feel instant.
The point is not a specific percentage. The point is that when volume is high and tasks are simple, the cost curve bends in favour of small models. We go deeper on this in our post about LLM cost optimisation with model routing and caching, which pairs naturally with an SLM strategy.
Consider a hypothetical Indian logistics startup that receives delivery exception messages from drivers in a mix of English, Hindi and Gujarati, typed quickly and full of shorthand. Every message needs to be tagged (address issue, customer unavailable, vehicle breakdown, damaged goods) and routed to the right ops team.
A first prototype sends every message to a large hosted model. It works well, but the team notices two problems. The bill grows linearly with delivery volume, and the response takes long enough that dispatchers start ignoring the auto-tags during busy hours.
The team collects a few thousand real messages, labels them, and fine-tunes a small multilingual model on the four tags. They keep the large model as a fallback for messages the small model is unsure about. In this kind of setup, the small model could handle most traffic quickly, while the large model only sees the odd, ambiguous cases. This is an illustrative scenario, not a reported client result, but it reflects the pattern we see teams reach when they measure before they commit.
Many providers offer small, cheap tiers of their model families. This is the lowest-effort path and a good place to start. You keep operational simplicity while testing whether a smaller tier meets your bar.
Running an open-weight model on your infrastructure gives control over data and predictable cost at scale. It also means you own uptime, scaling and updates. For regulated data this is often worth it.
For mobile and desktop features that must work offline or keep data local, quantized models can run directly on the device. The engineering effort is higher, since you manage model size, memory and battery, but the privacy and latency benefits are real.
Early-stage teams should rarely start by fine-tuning anything. Ship the first version on a capable hosted model, learn what users actually do, and log inputs and outputs responsibly. Once a task is stable and high volume, that log becomes the raw material for an evaluation set and, later, a small model. This sequencing keeps you fast at the start and efficient at scale.
If you are planning an AI feature and want help deciding between a large model, a small model or a hybrid, our team can walk through your use case as part of our AI development services.
The most common failure in model selection is a flattering test. Teams try ten hand-picked examples, see good results and declare victory. A better approach is to sample real traffic, include the messy cases and have more than one person label the answers. Where labels disagree, you have found a task definition problem, not a model problem.
Track more than accuracy. Measure latency at the slow end (the 95th percentile matters more than the average), cost per thousand calls and the share of outputs that fail format checks. For extraction tasks, count how often the JSON is invalid. For generation tasks, run a small human review each week to catch tone drift that metrics miss.
That last question is worth taking seriously. Not every problem needs a language model of any size. Sometimes a regular expression, a lookup table or a classic machine learning classifier is faster, cheaper and easier to explain to an auditor. Good AI development means choosing the simplest tool that works, then reaching for heavier tools only when the evidence demands it.
Small language models are not a replacement for frontier models. They are a sharper tool for a specific class of problems: narrow, repeatable, high-volume tasks where cost, speed and control matter. The winning approach is empirical. Define the task, build an honest evaluation set, compare a small model against your current baseline, and keep a fallback for hard cases. Teams that do this often find that the smallest model that clears the bar is also the one that makes the product feel fast and the unit economics work.