LangChain Launches smithtune CLI to Fine-Tune AI Agents
LangChain has launched LangSmith Fine-Tuning and the smithtune CLI in public beta, allowing developers to easily turn agent interaction trajectories into specialized, lower-cost models.

LangChain has released LangSmith Fine-Tuning and its accompanying command-line interface, smithtune, in public beta. The new toolset automates the pipeline of converting agent trajectories—the chronological sequences of messages, tool calls, and results from LangSmith traces—into datasets for supervised fine-tuning. By partnering with Fireworks and Baseten, LangChain allows developers to manage dataset curation, LoRA training, and evaluation directly from a single interface without needing to provision their own GPU infrastructure.
The smithtune workflow begins by pulling and filtering trajectories from LangSmith tracing projects. Developers can work with AI agents to label traces and filter out the best candidates for training based on a custom rubric. The tool then formats the data, splits it for training and validation, and submits the job to Fireworks managed SFT or Baseten Loops. After training, the CLI runs a replay evaluation to compare the fine-tuned model against its base version, generating a comparison link in LangSmith to analyze performance before deployment.
LangChain tested the new workflow on its own internal tools with notable results. In one test, the team fine-tuned the base Kimi K3 model for its Engine tool, which analyzes agent traces. The fine-tuned version outperformed both the base Kimi model and GPT-5.6 Sol on a subset of IssueBench, LangChain's internal benchmark for issue detection. In another trial using the OpenSWE Review tool, fine-tuning a Qwen-3.8-27B model raised its F1 score from 48.9% to 53.7% on an internal code-review evaluation set. Crucially, this performance boost was achieved while using 29.8% fewer model calls and 29.4% fewer tool requests.
The company notes that supervised fine-tuning is highly effective for applications performing repetitive tasks with established patterns of successful behavior. LangChain recommends that development teams first exhaust prompt and harness engineering before moving to fine-tuning. When recurring mistakes persist, curating high-quality trajectories with smithtune offers a clear path to reducing latency and operational costs.
This is our own summary of reporting by LangChain Blog



