Today we’re launching yoDEV Decisions, a free tool for yoDEV members that runs the same question on Jev and the LLM of your choice, row by row, and compares the results: accuracy, calibration, cost, and latency.
The idea came from our article on TypeSafe’s Jev. We concluded that a provider’s numbers don’t replace your own measurement, and that teams across Latin America need to measure with their own data, in their own language. Decisions is that measurement.
yoDEV covers calls to Jev. For the LLM side, you connect your own OpenRouter account.
What do the first results show?
For the public results, we chose four popular LLMs (GPT-4o mini, Gemini 2.5 Flash, Claude Haiku 4.5, and Llama 3.3 70B) and four public datasets, 200 rows each, with the question asked in Spanish. With your own data, you can compare Jev with any model available on OpenRouter, including free models.
| Dataset | Task | Jev | Best LLM |
|---|---|---|---|
| Banking77 | Intent of a bank customer, among 77 options | 82.0 % | 84.0 % (Gemini 2.5 Flash) |
| SMS spam | Is it spam? | 97.0 % | 98.0 % (Gemini 2.5 Flash) |
| Moderation in Spanish | Is it offensive? | 98.0 % | 98.5 % (GPT-4o mini and Llama 3.3 70B) |
| Sentiment in Spanish | Positive, neutral, or negative | 66.0 % | 75.0 % (Gemini 2.5 Flash) |
Three observations:
- Latency: Jev responded with a median of about 250 ms across all four datasets. The four LLMs took between 1 and 2.3 seconds.
- Cost: Jev was cheapest across all four datasets. On Banking77 it cost about $0.08 per 1,000 rows, versus $0.28 for Gemini 2.5 Flash and $1.84 for Claude Haiku 4.5.
- Accuracy: Jev doesn’t win on everything. On intent classification, spam, and moderation, it’s within two points or less of the best LLM, and even ahead of some. On Spanish sentiment, it trails by nine points.
That last result is the reason this tool exists: an average doesn’t tell you how a model performs on your task.
What if you combine Jev with an LLM?
Decisions also calculates a combined strategy: Jev responds when its confidence exceeds a threshold, and the remaining rows go to the LLM. The tool recommends that threshold based on your data.
On Spanish sentiment, that combination reached 74.5% accuracy by sending only 40.5% of rows to Gemini. You get nearly the LLM’s accuracy, with Jev handling most rows at its speed and cost.
How do I test it with my data?
- Log in to decisions.yodev.dev with your yoDEV account. If you don’t have one yet, registration is free.
- Connect OpenRouter for the LLM side and choose any of its models. You pay for those calls with your own balance, or choose a free model. Your OpenRouter key is saved in your browser.
- Choose a predefined dataset or upload a CSV with a
textcolumn and alabelcolumn with the expected answer: up to 2,000 rows. - Review the results: accuracy, calibration, cost, latency, the rows where models disagree, and the recommended combined strategy. Each run is automatically saved to your yoDEV Workplace.
You can ask the question in English and Spanish to measure the difference. Jev was trained primarily in English.
What happens to the data I upload?
Only metrics and row numbers are saved to the Workplace, never the text of your rows. Rows you upload are deleted 30 days after your last run with that dataset, and you can delete them earlier anytime.
What about OpenAI’s Decisions API?
On September 29, 2026, OpenAI announced at its DevDay a Decisions API: a model that chooses from developer-defined answers instead of generating text, the same category as Jev. As of September 30, it’s in limited preview, with no public API reference, pricing, or limits. 2026 DevDay Recap
That two providers are betting on the same idea confirms that bounded decisions are becoming their own piece of AI-powered applications. It also makes the age-old question more important: which works better on your task? When that API gets public access, it makes sense to measure it with the same method.
What are the limitations of these measurements?
- They’re 200 rows per dataset and a single run per model, conducted on September 29-30, 2026 with Jev 1.13.
- Latency depends on where you measure from. Test from your own region before drawing conclusions.
- Banking77 and SMS spam accuracy are measured with English texts; moderation and sentiment with Spanish texts.
- The moderation dataset contains offensive language.
Each results page links to the source and license of its dataset.
yoDEV Decisions is an independent yoDEV tool, with no affiliation with TypeSafe AI or OpenAI. Jev is a product of TypeSafe AI.
