AI for the work of running a business.Explore the benchmarks
itsryan.ai

Benchmarks

AI should help with the work of running a small business: estimates, customer replies, inventory, supplier bills, and everyday decisions. These tasks are informed by documented business workflows.

Compare published invoice results from three open-weight model configurations. Our own small-business evaluations are still in preparation. Startup tasks remain separate.

8 source-informed tasks across 4 business types. Workflow sources do not establish AI accuracy or time savings.

Published invoice results

Can an AI get the invoice total right? Three open-weight model configurations, tested on the same 200 synthetic invoices.

Published by jngb-labs / InvoiceBenchmark. 1,200 records across two methods. Runs dated April 20–21, 2026. Not run by ItsRyan.

The model reads the invoice and returns a total. Exact totals divided by all 200 attempts. Unparseable responses count as failures.

Model only · exact invoice totals
Model configurationExact totalsIncorrectUnparseableSource
Llama 3.3 70B Instruct77.0%154 / 200442CSV
Qwen3 8B51.5%103 / 2001978CSV
QwQ 32B51.0%102 / 2003761CSV

Relevant to invoice checking, not proof of small-business readiness. These are synthetic European-format text invoices, including large amounts; they do not test scans, your accounting system, or typical US small-business work.

Inspect the invoice results
Llama 3.3 70B Instruct · Model only
InvoiceExpected totalReported totalOutcome
INV-2026-0001€248,054.76€248,054.76Exact total
INV-2026-0002€183,313.83€183,213.83Incorrect total
INV-2026-0003€99,019.20€99,019.20Exact total
INV-2026-0004€222,953.88€222,953.88Exact total
INV-2026-0005€70,671.98€70,671.98Exact total
INV-2026-0006€14,856.50€14,856.50Exact total
INV-2026-0007€132,230.36€132,230.36Exact total
INV-2026-0008€201,192.60€201,192.60Exact total
INV-2026-0009€112,513.44€112,513.44Exact total
INV-2026-0010€229,496.31€229,496.31Exact total
1–10 of 200 invoices
Sources, scoring & limitations

Counts are recomputed from the publisher’s CSV files using exact decimal equality against the supplied answer key. All 1,200 row-level match flags agree with that calculation. This checks the published records, not the correctness of the answer key or a new model run.

We include all 200 attempts per model and method, including parse failures. Rates can therefore differ from the publisher’s summary of successfully parsed answers. Each method uses the same 200 invoice IDs; 1,200 records does not mean 1,200 distinct invoices.

These are selected published configurations, not a comprehensive leaderboard. Prompt, runtime, quantization, and hardware differences are not controlled here. No cost, speed, general capability, or statistical superiority claims are made. Rerun files are not combined with these original files.

Run dates come from CSV timestamps; the dataset description refers to May 2026. The dataset card declares MIT licensing. Snapshot captured 2026-10-01; revision f0699d8c94468fb2fefaa22f669684b22e68c3d3. Downloads include individual normalized results, source URLs, and SHA-256 hashes. Raw response text is omitted.

Read the pinned dataset documentation

Three focused comparisons

Small-business work, public research, and inspectable tests. These ItsRyan pilots are being prepared; we have not run our own evaluations yet.

Invoice arithmetic

Does this bill actually add up?

Check totals, discounts, and discrepancies before a bill reaches the payment queue.

Not yet evaluated

InvoiceBenchmark

5 synthetic examples inspected

Test scope & evidence

What we will measure

  • Exact monetary totals
  • Discrepancy detection
  • Valid structured output

Five proposed cases

  • Correct total
  • Wrong subtotal
  • Wrong final total
  • Conditional discount
  • Credit or negative amount

Limits of the evidence

European synthetic invoices, not typical US bills. They do not include purchase orders or delivery receipts, so they cannot validate supplier reconciliation.

Before evaluation

Audit a balanced case set and its tax and discount assumptions. The five inspected examples are a smoke sample, not the final benchmark.

Declared license: MIT

Retail sales analysis

What does this sales export tell us?

Summarize transactions and top products without losing cancellations or silently dropping messy records.

Not yet evaluated

UCI Online Retail

Historical data source identified

Test scope & evidence

What we will measure

  • Correct totals and rankings
  • Cancellation handling
  • Explicit exclusions and limits

Five proposed cases

  • Clean totals
  • Cancellations
  • Missing identifiers
  • Zero or negative values
  • Time-window product ranking

Limits of the evidence

Historical UK transactions, not a representative sample of current US businesses. Sales data alone cannot establish profit, stock availability, reorder quantities, or advertising ROI.

Before evaluation

Download and audit the dataset, freeze five transaction slices, and independently compute their answer keys.

Declared license: CC-BY-4.0

Policy-grounded replies

Can it answer without inventing an exception?

Draft customer replies from a written store policy and order facts, with clarification when information is missing.

Not yet evaluated

Sierra retail benchmark

Policy and task references inspected

Test scope & evidence

What we will measure

  • Policy compliance
  • Unsupported claims
  • Clarification and escalation

Five proposed cases

  • Eligible request
  • Ineligible request
  • Missing facts
  • Conflicting facts
  • Request to ignore the policy

Limits of the evidence

A proposed reply-only adaptation of a simulated, tool-using benchmark. It is not a reproduction of tau-bench and cannot inherit its published scores.

Before evaluation

Write five fresh cases with a versioned policy and obtain business-owner review. No live refunds or customer messages.

Declared license: MIT

Open-weight model shortlist

Two proposed baselines, not presumed winners. Local hardware fit, operating cost, and performance are still untested.

Proposed pilot: 15 distinct cases, 2 models, 3 repetitions. 90 runs planned; 0 performed.

  • Qwen3-8B Not evaluated

    Pilot baseline / Apache-2.0

    Text-only; record non-thinking mode, runtime, and quantization.

  • Pilot baseline / MIT

    Compact text baseline; runtime compatibility and local performance need testing.

  • Later candidate / Apache-2.0

    Larger baseline; check total weight memory before choosing a host.

Research library 8 sources

Publisher-declared licenses are recorded for review, not a blanket clearance to redistribute data. External benchmark scores are not part of the ItsRyan index.

InvoiceBenchmark

jngb-labs / Jakob Neugebauer / Pilot audit

Synthetic European invoices; not supplier delivery reconciliation. Audit the keys, monetary semantics, and notices before reuse.

MIT

UCI Online Retail

UCI / Daqing Chen / Pilot audit

Historical transactions, not stock, lead times, costs, or advertising data. Requires a new sales-analysis task ID.

CC-BY-4.0

Sierra retail benchmark

Sierra Research / Pilot reference

A tool-using simulated agent environment. Our reply-only adaptation cannot inherit its scores or benchmark name.

MIT

DABstep

Adyen / Hugging Face / Later research

Payment operations analytics, not advertising attribution. Some evaluation answers require the official grading flow.

CC-BY-4.0

Tokuhn small-merchant products

Tokuhn / Later research

Catalog grounding candidate, not live price or stock evidence. Card and viewer counts differ; deduplicate and review underlying content rights.

ODC-By

GDPval

OpenAI / Rights review needed

Public access does not establish redistribution rights. No license field or LICENSE file found in inspected Hub metadata; defer importing files.

License not established

Sources reviewed September 30, 2026. Dataset revisions are retained in the research plan. Startup research remains separate.

Business Work Index

Small-business tasks. Synthetic inputs. Inspectable rubrics and assumptions.

Your local evaluationsSelf-reported scores stored in this browser, not public rankings.

No evaluations for these tasks

No measured results have been recorded for this scope in this browser. Sources support the workflows, not model rankings.

Rankings & evaluation tools

Everyday business operations

Home services

From the first estimate to the final dispatch. Test whether a model can price a job, respect real constraints, and put the right person in the right place.

Sources & assumptions

Build a bathroom renovation estimate

An owner turns labor and material inputs into a customer quote.

  • Jobber: Quote Basics

    Service quotes can include quantities, unit prices, totals, scope, deposits, and customer messages.

Fixture assumptions: Rates, waste, contingency, and deposit are fixture assumptions, not market prices or recommended terms. Tax and permits are intentionally excluded.

Owner review: Check actual site conditions, prices, scope, taxes, permits, and terms before sending a quote.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

Reschedule an urgent service call

A dispatcher fits urgent work around team qualifications and fixed appointments.

  • Jobber: Visits

    Visits have scheduled times, assigned teams, locations, and rescheduling notifications.

Fixture assumptions: Travel takes exactly 30 minutes in this fixture. The example 9-11 emergency slot is one valid solution; other fully feasible schedules should be accepted.

Owner review: Confirm travel, certification, availability, breaks, and customer acceptance before moving a real booking.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

RankModelScore

No complete evaluations yet.

Everyday business operations

Retail & ecommerce

The details behind a good customer experience. Evaluate returns, inventory decisions, and the trade-offs that keep a small shop moving.

Sources & assumptions

Resolve a mixed returns inbox

A shop applies its written policy consistently to customer requests.

Fixture assumptions: The 30-day return window, 7-day damage window, and shipping exclusion are invented store rules, not universal Shopify rules or legal advice.

Owner review: Check the actual policy, consumer requirements, order status, and evidence before issuing a refund or replacement.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

Prepare a weekly reorder sheet

A retailer reconciles demand, stock on hand, and incoming inventory before ordering.

Fixture assumptions: The 14-day target and safety stock are supplied assumptions, not a validated forecasting method. The fixture assumes incoming stock arrives within the planning period.

Owner review: Check lead times, seasonal demand, stock accuracy, supplier minimums, cash limits, and delivery dates.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

RankModelScore

No complete evaluations yet.

Everyday business operations

Agencies & studios

Good client work starts with clear judgment. Test scope, campaign analysis, and the ability to turn a messy brief into a useful next step.

Sources & assumptions

Turn a client brief into a scoped proposal

A small service business turns a brief into a bounded estimate with clear exclusions.

Fixture assumptions: Hours, hourly rate, budget, and revision limits are invented. The sources support estimating workflows, not these rates or the legal adequacy of a contract.

Owner review: Confirm scope, availability, asset dependencies, payment terms, and contract language with the client.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

Explain which campaign deserves another test

An owner compares ad spend with leads and customers before a small follow-up test.

Fixture assumptions: All campaign figures are synthetic. The $200 per customer is advertising-only cost, not fully loaded CAC. Revenue, attribution, and statistical confidence are unknown.

Owner review: Verify conversion tracking, sales quality, revenue, and attribution. Treat the next budget allocation as a test, not a proven winner.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

RankModelScore

No complete evaluations yet.

Everyday business operations

Food & hospitality

Run the numbers before the doors open. Evaluate supplier invoices, catering quotes, and the operational details that protect your margins.

Sources & assumptions

Catch errors in a supplier invoice

An owner checks billed quantities and prices against the order and delivery before payment.

Fixture assumptions: Products, prices, and delivery discrepancies are synthetic. This exercise excludes tax and fees; the source documents bill linkage, not this specific cafe transaction.

Owner review: Check the actual purchase order, delivery receipt, agreed prices, credits, tax, and supplier response before paying.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

Quote a catering order with constraints

An owner assembles an itemized estimate and confirms availability before accepting an order.

Fixture assumptions: Menu prices, vegan count, budget, delivery fee, and 48-hour notice rule are invented. The source supports estimates, not dietary safety or restaurant operating standards.

Owner review: Confirm kitchen capacity, ingredients and allergen handling, timing, tax, and customer requirements before acceptance.

Source checked 2026-09-29. This supports task relevance, not model performance. Owner validation is pending.

RankModelScore

No complete evaluations yet.

Startups

A separate draft suite for product support, activation, customer discovery, and runway. Not part of the small-business scores.

Explore startup tasks
Ryan Widgeon
AI that earns its place in your business.

Benchmarks and practical guidance by Ryan Widgeon.

Work with Ryan
ItsRyan.ai Small Business Work IndexStarter suite · Synthetic business inputs · v1

Model: task evidence

Self-reported evaluations. Answers and grades have not been independently verified.

Sign up for updates

Get my latest AI tips, useful tools, and updates on my live class about using AI in your business.

Include your country code. Saved only when you opt in below.

Free updates. Unsubscribe anytime.