Read as article
SMITH Teaches Small AI Agents to Build Their Own Tools
By @sharedot · · 7 pages
Appier's NeurIPS-accepted SMITH framework trains agents to create and use tools in one loop, letting a 4B model beat a 30B baseline.
What Happened
Appier (TSE: 4180) announced that its paper 'Joint Optimization of Tool Creation and Use for Large Language Model Agents' has been accepted at NeurIPS, which the company describes as 'the Olympics of AI.' The paper introduces SMITH (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that lets an AI agent build tools and learn to use them inside a single training loop. Each tool is continuously refined based on real-world problem-solving results, closing the loop of building, using, verifying, and refining. CEO Chih-Han Yu framed the idea by analogy to people: humans turn problem-solving experience into reusable tools so they never start from scratch, and AI agents are now evolving the same way, according to the company's announcement.
Why It Is Surprising
The surprising part is the upset at model scale. According to PR Newswire, a roughly 4-billion-parameter model trained with SMITH built tools that outperformed every other method in the study on unseen tasks — including a baseline in which a roughly 30-billion-parameter model built tools on the fly. In most agentic AI setups, tool creation and tool use are assigned to separate models, so the tool-builder gets almost no feedback on whether its creations are clearly described, work reliably, or can be called correctly by other models. SMITH makes clarity and callability direct training feedback: Chieh-Yen Lin, a research scientist at Appier, said the model sees only each tool's description and parameter specifications, never the underlying code, so vague descriptions and failed calls feed straight back into learning.
The Evidence
Per the PR Newswire release, the training regimen starts from just 4 simple examples, from which the agent learns a method and builds tools, and is then tested on 16 harder, previously unseen problems. Only tools that solve new problems are kept and added to a shared tool library that multiple agents can draw from, with better tools replacing weaker ones over time. The study's reported findings: the small model's tools worked even for a lightweight model of only about 350 million parameters, and also boosted larger models, enabling flexible division of labor. Most striking, repeated step-by-step reasoning was converted into callable tools, cutting average output from 3,206 tokens to about 100 — roughly a 32-fold efficiency gain — while maintaining task performance.
The Stakes
For builders, this challenges the default assumption that effective tool creation must depend on bigger models. Appier positions the work as a path to scalable multi-agent collaboration: proven tools can be shared across models of all sizes and across tasks, reducing the token cost of repeated reasoning. The company points to repetitive enterprise operations — converting financial metrics, processing data, querying reports, checking rules, and routing customer service cases — as candidates for turning scattered methods into shared, verifiable tools. It also sees applications in advertising and marketing, where agents handling customer data, personalization, and ad buying could share proven tools and consistent business rules. For practitioners, the development pairs with broader 2026 agent-building practice: as the Bitcoin Foundation's guide notes, reliable agents combine models, tools, memory, workflows, guardrails, and monitoring into one system.
What Comes Next
Appier says it will continue advancing Agentic AI through research and bringing it into real-world applications with measurable business results, and Lin indicated the team hopes to build models that continuously interact with their environment and take on a wider range of tasks. For builders watching the space, the practical takeaway is a shift from engineers pre-building and reconfiguring fixed APIs toward agents that maintain and share living tool libraries. The Bitcoin Foundation's practitioner guide notes that API descriptions must stay precise and that models need to understand what each function does and which parameters it requires — precisely the failure modes SMITH turns into training signal.