A small language model can produce a valid-looking function call without being reliable enough to run an AI agent. The harder test is whether it selects the correct tool, supplies schema-valid arguments, recognizes when no tool applies, and continues after receiving a result.
Compact models can run closer to the user, reduce latency and hardware requirements, and keep some requests off external APIs. Their limits appear when tools have similar names, parameters are missing, or a workflow spans several steps.
The ten models below are matched to workloads ranging from narrow mobile actions to multimodal and reasoning-heavy local agents.
Function Calling Is Only One Part of an AI Agent
Function calling connects a language model to external software. The model receives function descriptions and returns a structured request identifying a function and its arguments. Application code must validate the request, execute it, return the result, and decide whether another model turn is required. Google’s FunctionGemma workflow, for example, separates the model-generated call, developer-side execution, and final model response into distinct stages.
A weather lookup may need one call. A purchasing agent may need to search inventory, compare results, request approval, and create an order. That requires state management across several calls.
Three capability levels are useful:
- Format-compatible: Produces structured calls with the correct template.
- Tool-specialized: Trained to select functions and generate arguments.
- Agent-oriented: Uses tool results across multiple steps or turns.
Sensitive actions still require application-side permissions, validation, and confirmation.
How the Models Were Selected
The list focuses on models of roughly four billion parameters or fewer with officially documented tool support. Selection also considered licensing, local inference, context capacity, irrelevant-tool handling, and whether support extends beyond single calls.
Official documentation establishes intended formats and uses, not guaranteed performance on every API. The final test must use the actual tools, prompts, and failure cases the agent will encounter.
| Model | Size | Best fit | Main limitation |
|---|---|---|---|
| FunctionGemma | 270M | Fine-tuned device actions | Needs specialization |
| xLAM-2 | 1B | Agent research | Non-commercial license |
| LFM2-Tool | 1.2B | Low-latency edge calls | Custom license |
| Hammer 2.1 | 1.5B | Multi-step workflows | Non-commercial license |
| Qwen3 | 1.7B | MCP and local assistants | General rather than specialized |
| Llama 3.2 | 3B | Broad runtime support | Older generation |
| SmolLM3 | 3B | Custom self-hosted stacks | Tool format needs care |
| Granite 4.1 | 3B | Enterprise automation | Smaller ecosystem |
| Ministral 3 | 3B | Multimodal agents | Long histories need testing |
| Phi-4 Mini | 3.8B | Reasoning-heavy assistants | Larger footprint |
FunctionGemma 270M
FunctionGemma is a function-calling version of Gemma 3 270M. Google describes it as a foundation for custom local action models, not a general dialogue assistant. It has a 32K context and is intended for task-specific fine-tuning, including multi-turn use cases.
It suits small, stable action sets such as mobile controls, smart-home commands, or application functions. Google also proposes using it as a local router that escalates complex requests to a larger model. Its size is attractive, but developers should expect to prepare representative training data rather than depend on zero-shot prompting.
xLAM-2-1B-fc-r
Salesforce’s xLAM-2-1B-fc-r is fine-tuned for function calling and multi-turn conversations. It has a 32K default context, supports Transformers and vLLM, and has official GGUF quantizations.
It is more relevant than a general 1B chat model for experimenting with compact action models and extended tool interactions. The “r” suffix marks it as a research release, and its CC BY-NC 4.0 license prevents unrestricted commercial deployment.
LFM2-1.2B-Tool
Liquid AI designed LFM2-1.2B-Tool as a concise, non-thinking tool caller for API calls, database queries, vehicles, IoT devices, and other latency-sensitive environments.
Avoiding explicit reasoning can help when requests map directly to known functions and additional reasoning tokens only add delay. The model has deployment examples for Transformers and SGLang, plus local quantizations. It uses Liquid AI’s LFM 1.0 license, which teams should review before commercial use.
Hammer 2.1 1.5B
Hammer 2.1 is fine-tuned from Qwen2.5-Coder and supports single-turn, multi-step, and multi-turn function calling. Its training also targets irrelevant-tool detection, allowing it to answer normally when none of the available functions fit.
The 1.5B model supports Hugging Face tool templates, vLLM deployment, and a 32K context. It is useful for evaluating workflows involving intermediate calls, but its CC BY-NC 4.0 license limits commercial use.
Qwen3-1.7B
Qwen3-1.7B has a 32,768-token context and can switch between thinking and non-thinking modes. Qwen recommends Qwen-Agent for tool templates, parsing, integrated tools, and Model Context Protocol configurations. It is released under Apache 2.0.
This model fits assistants that must converse and explain as well as call tools. Non-thinking mode can handle direct requests, while thinking mode can be reserved for planning. A specialist fine-tuned on one API surface may still be more predictable.
Llama 3.2 3B Instruct
Llama 3.2 3B Instruct remains useful because of its broad runtime and framework support. Meta documents a 128K context, agentic use cases, and evaluations on BFCL V2 and Nexus.
Meta’s reported BFCL V2 accuracy was 67.0 for the 3B BF16 model, compared with 25.7 for the 1B BF16 model. Treat those figures as evidence for choosing the 3B variant, not as a direct comparison with results from other benchmark versions. Llama 3.2 uses Meta’s custom community license.
SmolLM3-3B
SmolLM3 is an Apache 2.0 model with published training details. It was trained at 64K context and supports extension to 128K through YaRN. Its chat template supports reasoning and non-reasoning modes plus XML-described or Python-style tools.
It works with Transformers and vLLM’s Hermes tool-call parser, making it suitable for teams that want control over prompting, parsing, quantization, and fine-tuning. It standardizes one call format and validates every generated argument.
Granite 4.1 3B
IBM’s Granite 4.1 3B is intended for edge and resource-constrained deployments. IBM reports improvements in tool calling, instruction following, chat, coding, and mathematical reasoning from its updated post-training pipeline. The model is available under Apache 2.0.
Granite accepts tools through an OpenAI-style function-definition schema, which helps teams already using that interface. It is a sensible candidate for internal business workflows, although its third-party agent ecosystem is smaller than Llama’s or Qwen’s.
Ministral 3 3B
Ministral 3 3B combines text and vision capabilities in an edge-oriented model. Mistral documents a 256K context, Apache 2.0 weights, function calling, structured outputs, built-in tools, and integration with its Agents and Conversations APIs.
It fits agents that must inspect an image or document before choosing an action. The advertised context is capacity, not proof of reliable reasoning across an entire long history, so document-heavy and extended conversations still need testing.
Phi-4-mini-instruct
Phi-4-mini-instruct has about 3.8B parameters and a 128K context. Microsoft provides a tool-enabled prompt format in which function definitions are supplied as JSON inside dedicated tool tokens. The model uses the MIT license.
Its reasoning-focused training makes it worth testing when instructions must be interpreted before an action is selected. The trade-off is a larger memory and latency footprint than sub-2B specialists, which may be unnecessary for repetitive commands.
Choose According to Workload and Failure Cost
For a narrow action set, start with FunctionGemma and plan for fine-tuning. LFM2-Tool targets fast direct calls, while Qwen3-1.7B is a balanced commercial option below 2B. SmolLM3, Granite, Ministral, Llama, and Phi add different combinations of openness, enterprise focus, vision, ecosystem support, context, and reasoning.
The consequences of failure should determine safeguards. A mistaken weather lookup is inconvenient; an incorrect payment, database deletion, or infrastructure command can be costly. High-risk tools need permissions, deterministic schema validation, argument constraints, audit logs, and human confirmation.
A practical architecture lets a small model process common, low-risk requests and escalates ambiguous or sensitive cases.
Test the Model Against Your Own Functions
Build an evaluation set from real schemas and user language. Include similar tool names, missing parameters, invalid types, requests requiring no tool, parallel calls, dependencies on earlier results, execution failures, and unauthorized actions.
Measure correct tool selection, argument accuracy, false calls, recovery after errors, latency, and performance as history grows. Run each case repeatedly because one successful response does not establish production reliability.
