Creating The AI Model You Wish Already Existed
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Creating The AI Model You Wish Already Existed on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A Hugging Face contributor reports using the ML-intern agent to build and publish seven custom models over several days. Detailed examples include a small prompt rewriter and a citrus image classifier, but the performance and cost figures are self-reported and have not been independently verified.

A Hugging Face contributor says they used the platform’s ML-intern agent to build and publish seven custom models over several days, including a small prompt rewriter and a citrus disease image classifier, as described in the original analysis. The account offers a practical example of agent-assisted model development, but its performance and compute-cost figures are self-reported, not independent evaluations of the agent or a guarantee of results for other users.

The contributor said the first project addressed a specific limitation: the prompt rewriter included with Qwen-Image 2.1 was too large for their needs. They described the official model as having 9 billion parameters and requiring about 20 GB of memory. After finding compressed versions of the large model but no smaller alternative, they trained a 0.8-billion-parameter rewriter using outputs from the larger model as examples.

According to the contributor, the smaller model produced valid output in 99.7% of cases and used about one-quarter as many tokens as its teacher. They put the compute cost, including labeling 8,797 example requests with the larger model, at about $16. The account does not specify how validity was measured or provide an independent assessment of the comparison.

Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies in images. The contributor said the training set had 3,017 annotated images across 21 categories. On 335 test photos, they reported that the base model selected the correct problem 14.9% of the time, compared with 52.8% for the fine-tuned model after two epochs on one A10G GPU. The stated compute cost was about $1.90.

The contributor also described a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the latter, the agent generated 24,722 transparent images of scanned household objects from 24 angles, with some objects reserved for testing. Training took about 90 minutes on one A100; the contributor estimated total compute at about $16, including failed jobs that were resubmitted. The account gives detailed figures for only some of the seven projects.

At a glance
reportWhen: Reported in an account on ThorstenMeyer…
The developmentA Hugging Face contributor has described using the ML-intern agent to create and publish seven custom models, sharing selected costs, test results and workflow details.
At a glance
reportWhen: Reported last week; the projects were b…
The developmentA Hugging Face contributor says an AI agent called ML-intern helped plan, train, evaluate and publish seven custom models on the Hub over several days.

Lowering Barriers to Model Customization

The report matters because it shows how an AI agent might reduce the hands-on work required to turn a narrow need into a trained and published model. The contributor said ML-intern proposed plans, requested approval before paid work, ran preliminary tests, then handled training, evaluation and publication using Hugging Face hardware. If similar workflows prove reliable across users and tasks, developers may be able to try specialized models without coordinating every technical step themselves.

The reported examples also show why cost figures need careful interpretation. The amounts describe compute spending for these projects, not a full accounting of data preparation, prompt writing, review time or other expenses. Nor do a few successful examples establish what a typical user would pay or how well the approach would work on a different task.

Evaluation against an untuned model can make a reported gain easier to interpret: the citrus figures compare both models on the same stated set of 335 photos. But test-set results alone do not show how a model will perform on new images or in real-world use. The author also reported that later character-LoRA checkpoints began influencing prompts unrelated to the intended character, illustrating a possible cost of continued training: a model may generalize the desired style too broadly.

Amazon

AI model training GPU

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Contributor Ran Each Project

The contributor said each project began as a message in HuggingChat with ML-intern enabled. The agent proposed a plan and asked for a spending limit before paid work. When a prompt did not include a budget, it offered options and asked the user to choose. The reported workflow then included a small test run before the main training job, followed by evaluation and publication.

The contributor said their instructions grew more detailed over time, from about 450 words for the first project to nearly 2,000 for the sixth. They included the dataset, base model and training script, and asked for a baseline, a small smoke test and a spending cap. One instruction called for the base model’s zero-shot score on the same metric before training, to make the gain visible. The contributor says all seven prompts are available in a public GitHub repository, while model cards and evaluations were published on Hugging Face.

The source presents these steps as one user’s account of directing the agent. It does not amount to an independent test of ML-intern, and the supplied material does not provide full descriptions of all seven projects. The author’s prompt practices offer context for the reported results, but do not establish that the same process will produce comparable outcomes elsewhere.

““Also report the base model’s zero-shot score on the same metric before training so we can see the gain.””

— The Hugging Face contributor

Amazon

machine learning model development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Checks Are Still Missing

The figures in the account have not been independently replicated in the supplied material. It does not provide detailed evaluation protocols for every model, comparable results from other users, or full descriptions of all seven projects. The contributor’s reported 99.7% valid-output rate is not accompanied by a definition of valid output or an explanation of how the rate was calculated.

Other open questions include whether test images were independently reviewed, how data quality was checked, and how well the models perform on examples outside the reported test sets. The stated costs cover compute according to the contributor, but the account does not provide a complete breakdown of labor or all project expenses. Those limits make it difficult to determine whether the reported results and costs are typical.

Amazon

AI image classifier training software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

More Projects Could Test the Claims

The contributor says the models, evaluations and project prompts are available through Hugging Face and GitHub, giving readers material to inspect. Further comparisons across users, datasets and tasks would help show whether the reported performance gains and compute costs recur beyond these examples.

The account’s suggested practices for future work include establishing a baseline, running a small test and agreeing on a spending limit before training. Whether those checks are sufficient will depend on each task and its risks. No independent follow-up results or next evaluation date are specified in the supplied material.

Amazon

custom AI model creation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did the contributor use ML-intern to build?

The contributor says the agent helped build and publish seven custom models over several days. The supplied account gives detailed examples for a 0.8-billion-parameter prompt rewriter, a citrus image classifier and two LoRA projects.

How much did the projects cost?

The contributor reported about $1.90 in compute for the citrus classifier and about $16 for each of two other detailed projects. These are self-reported compute costs, not complete project budgets.

Are the reported results independently verified?

Not in the supplied account. The performance figures come from the contributor, and the material does not report independent replication or provide detailed evaluation methods for every model.

What did the citrus model classify?

It was fine-tuned to identify citrus pests, diseases and nutritional deficiencies from images. The contributor reported a higher correct-identification rate than the base model on a 335-photo test set, but broader performance remains unclear.

Where can readers inspect the work?

The contributor says the models and evaluations are available on Hugging Face, and that all seven project prompts are in a public GitHub repository. The supplied material does not include the repository address.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

WordPress Surges In Global Coverage

WordPress experiences a surge in international media coverage, with 34 mentions in recent monitoring, highlighting its growing influence.

OpenAI Dev Day 2026: Key News For Developers And Builders

OpenAI announced agents, plugin upgrades, new APIs and models at DevDay 2026, including Dots and a limited-preview Decisions API.

Vaneck Fabless Semiconductor Surges In Global Coverage

Vaneck Fabless Semiconductor ETF experiences a surge in international coverage, with 13 mentions in recent media analysis, indicating rising global interest.

The AI Search That Discovered Buried Data

An AI model demonstrated the ability to locate critical hidden data within company files, impacting sales and trustworthiness in AI automation.