🔍 Read the full analysis: 24 Ways To Use Jev From AI Planning To Decision-Making on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer’s Sept. 29 article maps 24 possible uses for Jev across publishing, commerce, software, business operations and the home. It says three uses are live, 12 are strong fits, seven need measurement and two are poor fits, while the available source details only the first six publishing examples.
Thorsten Meyer published a 24-use-case assessment of Jev on Sept. 29, describing applications across publishing, commerce, software, business operations and the home. Meyer says three uses are already live in his publishing operation, while 12 are strong fits, seven need measurement and two are poor fits; the source material available here gives details for only the first six publishing examples.
Meyer describes Jev as a system that receives text or JSON plus typed questions and returns answers that software can use to make decisions. It does not write, summarize or extract, he says. A call can carry several questions and takes about 0.3 to 0.9 seconds, at a stated cost of about $0.04 per million input tokens. Answer types include yes-or-no probabilities, choices with per-option probabilities and confidence, and scores on ordered levels.
The three reported live uses are a relevance gate for matching stories to a site, an English-language check, and a fallback topic classifier. Meyer says the language check processed 78,889 articles for $2.01, finding 1,576 non-English articles and fixing 1,553. The relevance gate assessed about 10,000 story-site pairings in three days, with 22% judged clearly on-topic. He reports 89% agreement between the fallback classifier and a frontier LLM overall, rising to 97%–99% for answers with confidence of at least 0.8.
Among the six publishing examples described, Meyer labels disclosure checking and comment moderation strong fits. Thin-source detection, product matching in roundups and headline-quality assessment need measurement first. Same-event deduplication is a poor fit in his canary: it found no duplicates, leaving no demonstrated problem for the tool to address.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Cheap Judgments Could Help
Meyer proposes screening automation ideas by asking whether a system must make many small, narrow judgments, whether errors are inexpensive or uncertain cases can be referred elsewhere, and whether an existing heuristic has a measured failure. He says these criteria can help identify workflows for testing the tool.
His suggested pattern is to automate clear cases and route uncertain ones for more capable review or human attention. In the publishing examples, that means holding back doubtful disclosure checks rather than automatically publishing an unverified result. The potential benefit is broader coverage at low per-call cost; outcomes depend on the tool’s accuracy in each workflow and the cost of mistakes.
A Confidence-Based Workflow
Meyer says Jev returns structured answers with confidence, leaving the application code to decide what action follows. In his measurement on a 31-topic classification task, he reports 97%–99% agreement with a frontier LLM when Jev’s confidence was 0.8 or higher, compared with 42% agreement below 0.5. These are Meyer’s reported results for that task, not a general accuracy guarantee.
His proposed fit test has four parts: high volume, a narrow question, low-cost errors or a route for uncertain cases, and a visibly failing existing heuristic. He advises replaying 300–500 past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements before wiring a use into production. He says to proceed only where the high-confidence band reaches 95%, then use a separate feature flag, a small canary and a gradual rollout.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
What the Published Excerpt Covers
The source material supplied for this article ends partway through the commerce and customer-operations section. It does not include the remaining use cases, so the reported totals of 24 applications, 12 strong fits, seven requiring measurement and two poor fits cannot be independently matched against a complete list here.
The performance and cost figures are attributed to Meyer; the supplied material does not include independent testing details, datasets or a third-party evaluation. It also does not establish whether the results will generalize to other publishers or workflows. For examples tagged “measure first,” the source itself says the existing heuristic’s error rate has not yet been demonstrated.
Measure Before Wider Rollout
Meyer’s next step for a prospective use is to compare Jev with real past decisions, examine disagreements and confirm that its high-confidence results reach the stated threshold. He recommends enabling a successful use behind a separate flag, beginning with a 5%–10% canary before expanding deployment. The article excerpt does not provide dates for further releases or evaluations.
Key Questions
What does Jev do?
According to Meyer, Jev answers typed questions about supplied text or JSON with structured results, such as probabilities, classifications or scores. Application code determines what to do with those answers.
Which Jev uses does Meyer say are live?
He identifies a story-to-site relevance gate, an English-language check and a fallback topic classifier in his publishing operation.
How much did the reported article scan cost?
Meyer says a scan of 78,889 articles cost $2.01 and identified 1,576 non-English articles, of which 1,553 were fixed.
How should a team decide whether to use Jev?
Meyer’s test calls for high volume, a narrow question, low-cost errors or a path for uncertain cases, and evidence that the current heuristic fails. He recommends testing against 300–500 past decisions before rollout.
Are all 24 proposed applications detailed in the available material?
No. The supplied excerpt describes six publishing examples and begins a commerce section, but ends before the full list appears.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
