When No AI Model Fits, Build Your Own
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: When No AI Model Fits, Build Your Own on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A Hugging Face contributor reports using its ML-intern agent to build and publish seven custom models over several days, including a small prompt rewriter and a citrus-disease image model. The contributor reports low compute costs and performance gains in selected tests, but the figures are self-reported and have not been independently replicated.

A Hugging Face contributor says the platform’s ML-intern agent helped build and publish seven custom models over several days, including a compact prompt rewriter and an image model for citrus problems, as described in the original analysis. The account offers a practical example of agent-assisted model development, but its performance and cost figures are the contributor’s own reports, not an independent evaluation.

The first project addressed a 0.8-billion-parameter prompt rewriter for Qwen-Image 2.1. The contributor wanted a smaller alternative to the included 9-billion-parameter model, which they said needs about 20 GB of memory and can produce thousands of tokens before returning a paragraph. They reported that the compact model produced valid output 99.7% of the time and used about one-quarter as many tokens as its larger teacher. The reported compute cost was about $16, including having the teacher label 8,797 example requests.

Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies in photographs. The contributor said the dataset contained 3,017 annotated images across 21 categories. On 335 test photos, the base model identified the correct problem in 14.9% of cases, compared with 52.8% for the fine-tuned model after two training epochs on one A10G GPU. The reported compute cost was about $1.90.

The account also describes a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the camera project, the agent generated 24,722 transparent images of household objects from 24 angles. Training took about 90 minutes on one A100; the contributor put total compute spending at about $16, including failed jobs that had to be resubmitted. The source gives details for only some of the seven projects.

At a glance
reportWhen: Reported recently; the source account d…
The developmentA Hugging Face contributor has published an account of using the ML-intern agent to plan, train, evaluate and publish seven custom models.
At a glance
reportWhen: Reported last week; the projects were b…
The developmentA Hugging Face contributor says an AI agent called ML-intern helped plan, train, evaluate and publish seven custom models on the Hub over several days.

Lowering the Barrier to Custom Models

The account suggests that an agent can take on parts of a workflow that otherwise require a developer to coordinate data preparation, test runs, training, evaluation and publication. If this approach works reliably for other users, it could make it easier to try task-specific models rather than adapting a larger general-purpose model by hand. The examples are small, bounded projects, however, and do not establish that the same costs or results would apply to more demanding applications.

The reported citrus results also show why a comparison matters: the contributor measured the tuned model against the base model on the same test set. That makes the stated difference more informative than a score without a baseline, while leaving open questions about the test set’s quality and how well results carry over to new images. The contributor also reported that later character-LoRA checkpoints began affecting prompts unrelated to the character, illustrating that training can introduce unwanted behavior as well as improve a target task.

Costs are another useful but limited part of the example. The amounts cited cover compute spending, not a complete accounting of dataset creation, prompt writing, human review or other possible expenses. They should be read as reported costs for these projects, not as a price estimate for building a custom model generally.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Agent Workflow Was Used

The contributor says each project began as a request in HuggingChat with ML-intern enabled. The agent proposed a plan, asked the user to approve spending before paid work, ran a small test and then handled training, evaluation and publication using Hugging Face hardware. When a prompt did not include a budget, the agent reportedly offered options and asked the user to choose.

The prompts became more detailed across the projects, growing from about 450 words for the first to nearly 2,000 for the sixth. According to the contributor, they specified the dataset, base model and training script, and asked for checks such as a baseline result, a small test run and a spending cap. The contributor said the seven prompts are available in a public GitHub repository, while model cards and evaluations were published on Hugging Face.

For the prompt-rewriter project, the contributor first found compressed versions of the 9-billion-parameter model on the Hub but said they did not find a smaller alternative that met the need. The resulting model was distilled using examples labeled by the larger model. That project and the citrus classifier illustrate different uses of customization: producing shorter text and adapting an image model to a defined set of crop problems.

“Also report the base model’s zero-shot score on the same metric before training so we can see the gain.”

— The Hugging Face contributor

What the Report Does Not Establish

The figures are self-reported; the account does not provide independent replication or a comparable evaluation of all seven models. It also does not give full descriptions of every project. For the reported tests, the available material does not specify all evaluation procedures, how test images were checked, or how the 99.7% valid-output rate was calculated.

It remains unclear whether the results would hold on different datasets, unfamiliar images, other hardware or larger budgets. The compute estimates do not include a full breakdown of labor and other expenses, and the account does not establish how consistently the agent handles data quality or failed runs across users. Those limits mean the examples show what one contributor says they achieved, rather than typical outcomes for the platform.

Published Projects Invite Further Checks

The contributor says the models, evaluations and project prompts are available through Hugging Face and GitHub for readers to inspect. Independent tests using the published models and clearly described datasets could help assess whether the reported gains hold beyond the contributor’s own runs.

Comparisons across users, tasks and test sets would also show whether the reported compute costs are typical. The contributor’s stated process—setting a budget, measuring a baseline and running a small test before full training—provides details readers can check against future projects. No broader replication results or further milestones are provided in the account.

Key Questions

What is ML-intern?

In the account, ML-intern is an agent used through HuggingChat to plan model work, request spending approval, run tests, train models, evaluate them and publish results on Hugging Face.

How many models did the contributor report building?

The contributor said they built and published seven models over several days. The source provides detailed figures for several projects, not all seven.

What results were reported for the citrus model?

On 335 test photos, the contributor reported that the fine-tuned model identified the correct citrus problem 52.8% of the time, compared with 14.9% for the base model. The test set covered a project trained on 3,017 annotated images across 21 categories.

Are the reported costs and results independently verified?

No independent replication is described. The performance figures and compute costs come from the contributor’s account, and the costs do not represent a complete accounting of time or every possible expense.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Virtual Reality 2.0: Beyond Gaming and Entertainment

Beyond gaming, Virtual Reality 2.0 is revolutionizing industries with immersive, interactive experiences that could change the way we learn, work, and innovate—discover how.

Autonomous Vehicles and Drones: Where Are We Now?

Discover how autonomous vehicles and drones are evolving with cutting-edge sensors and regulatory challenges shaping their future.

AI in Finance: Fraud Detection vs Forecasting—Different Beasts

AIThis post was created with the assistance of artificial intelligence (AI).AI in…

The AI That Reads the Footnotes May Be the One That Wins the Contract

A €55,000 contract exposed a crucial gap between AI agents: all saw the opportunity, but only those that read the buried file won the deal.