Daily Podcast full article
Appier agents build their own tools
Appier’s freshly announced NeurIPS research reframes agent autonomy: the next jump is not just connecting an AI system to more APIs, but training it to create, validate, reuse and share tools. The promise is cheaper, more adaptable enterprise workflows; the hard part is proving that self-built tools stay secure, understandable and durable when they leave the lab.

What happened
Appier announced that its paper, “Joint Optimization of Tool Creation and Use for Large Language Model Agents,” has been accepted at NeurIPS, introducing SMITH, short for Schema-grounded Multi-task Iterative Tool Honing . The core claim is simple but important: agents should not only call tools that engineers prepared in advance; they should learn to build tools, test them, keep the ones that generalize and reuse them across tasks .
The company’s announcement describes SMITH as a reinforcement learning framework that trains tool creation and tool use inside one loop, so the same learning process is judged on whether a generated tool can actually solve later problems . That is the meaningful shift. Many agent systems already look impressive when they can choose among existing functions, search endpoints, calculators or internal APIs. Appier is arguing that the more scalable path is to let agents turn repeated reasoning patterns into callable, shareable tools.
Taiwan’s Economic Daily framed the research as a move from “using tools” toward “autonomously creating and optimizing tools,” noting that SMITH combines the ability to build and correctly use tools in response to the common failure mode where an AI can generate a tool but not reliably use it . Appier’s Japanese release makes the same point: the work targets the gap between tool generation and effective tool operation, rather than treating code creation as a standalone trick .
Why SMITH matters
The practical problem is that enterprise agents usually depend on integrations made by people. If a workflow changes, if a data source moves, or if a business rule is rewritten, engineers often have to update the available toolset. Appier’s release explicitly identifies that dependency: many systems still rely on developers to prebuild APIs or fixed tools, and those tools may need to be rebuilt when data sources, tasks or business needs change .
SMITH tries to close that gap by training an agent to infer a procedure from examples, express it as a tool, and then prove the tool works on harder, unseen problems. Appier says the build side starts from four simple examples, after which the generated tool is tested on sixteen harder unseen questions; only tools that pass are kept in a shared pool . That design is important because it treats “tool creation” less like a coding demo and more like a quality-control pipeline.
The “schema-grounded” part is also central. In the use stage, the model sees the tool description and parameter specification rather than the underlying code, so a vague description or badly designed argument list becomes a direct training failure . In other words, the agent is not only rewarded for writing code that runs; it is rewarded for creating an interface that another agent, or a later version of itself, can understand and call.
The numbers Appier is highlighting
Appier says a roughly 4-billion-parameter model trained with SMITH outperformed other methods in the study on unseen tool-creation tasks and even beat a baseline where a roughly 30-billion-parameter model created tools on the fly . Economic Daily reported the same comparison, emphasizing that effective tool building may not require simply scaling to a larger model .
The efficiency claim is just as striking. According to Appier, SMITH reduced average output from 3,206 tokens under conventional step-by-step reasoning to about 100 tokens by converting repeated reasoning into reusable tools, a roughly 32-fold reduction . Appier’s Chinese-language release also highlights the same drop from 3,206 tokens to about 100 tokens, presenting it as a route to lower inference cost while maintaining task performance .
The research also argues that tools generated by a smaller trained model can travel across model sizes. Appier says tools produced by the small model still handled new tasks when used by a lightweight model of about 350 million parameters, and also improved larger models’ performance . That matters for multi-agent systems because not every agent in a workflow should need to be the biggest or most expensive model. A smaller specialist that can create a reusable tool for other agents changes the economics of orchestration.
From clever benchmark to business workflow
The obvious business reading is that Appier wants agentic AI to become cumulative. In today’s agent stacks, too much work is ephemeral: a model reasons through a task, emits an answer and discards the intermediate method. SMITH points toward a different pattern, where successful reasoning is compressed into an executable skill that can be inspected, reused and improved.
Appier connects that idea to enterprise operations such as financial metric conversion, data processing, report queries, rule checks and customer-service routing . In marketing and advertising, the company says agents working on customer data, personalization, service and ad buying could share verified tools and consistent business rules . That is consistent with Appier’s broader positioning as an AI-native company focused on AdTech and MarTech, with products in advertising, personalization and data clouds .
The commercial appeal is clear. If an agent can learn a business procedure from a few examples and turn it into a reusable tool, teams might reduce repeated integration work. Instead of asking a model to reason from scratch every time a similar case appears, the workflow could call a verified tool. That is faster, cheaper and, in principle, easier to monitor.
The security question Appier now has to answer
The risk is also clear: self-built tools are still software. They can contain bugs, ambiguous parameters, hidden assumptions and unsafe side effects. Appier’s own description of SMITH acknowledges the failure modes indirectly by emphasizing tool descriptions, parameter specifications, execution failures and validation before a tool enters the shared library . That is the right vocabulary for a research system, but production deployment will require more than benchmark validation.
The key enterprise question is not whether an agent can make a clever Python function. It is whether the resulting tool can be permissioned, logged, reviewed, versioned and revoked. A self-built tool that only transforms a table is one risk category. A self-built tool that touches customer records, spends ad budget, updates a CRM field or triggers an external service is another. The moment generated tools become reusable infrastructure, they need the same governance as human-written integrations.
Interpretability will matter too. SMITH’s use-stage discipline, where models rely on descriptions and parameter specifications rather than code visibility, could encourage clearer interfaces . But enterprises will still want to know what the code does, where it came from, which examples shaped it, which tests it passed and which agent last modified it. Otherwise, a tool library can become a junk drawer of brittle one-off scripts.
The bigger signal
Appier’s NeurIPS acceptance is best read as a research signal rather than a finished enterprise product announcement. Still, it lands at the right pressure point for agents. The bottleneck is shifting from “Can the model call a tool?” to “Can the system continuously acquire useful capabilities without multiplying operational risk?”
If SMITH’s results hold up outside controlled tasks, the next generation of agents may look less like chatbots with plug-ins and more like apprentices that turn repeated work into shared procedures. That would make agent systems more adaptable, especially when workflows are unfamiliar or under-documented. It would also force companies to build new controls around tool creation itself: approval queues, sandboxing, dependency scanning, provenance records, permission scopes and rollback.
For now, Appier has shown the research direction clearly: agents are entering crafting mode. The useful version is not an agent that writes random code whenever it gets stuck. It is an agent that builds tools, proves they work, explains how to call them, shares them safely and forgets the bad ones. That is where the autonomy story becomes commercially interesting, and where the permissions menu becomes the real product.
Sources from the last 72 hours
- [1]Appier Research Accepted at NeurIPS: AI Agents Learn Not Only to Use Tools, but to Build Their OwnSep 30, 2026, 3:06 PM
- [2]Appier Research Accepted at NeurIPS: AI Agents Learn Not Only to Use Tools, but to Build Their OwnSep 30, 2026, 2:00 AM
- [3]Appier SMITH 登 NeurIPS:AI Agent 能自主打造與優化工具Sep 29, 2026, 11:54 AM
- [4]Appierの研究論文がNeurIPSに採択AIエージェントはツールを「使う」だけでなく「自ら作る」段階へSep 30, 2026, 5:43 AM
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.

Comments
Be the first to comment.