AINNA Research

Operational Benchmark

NeuralOps Token Efficiency Benchmark

An AINNA operational benchmark recorded approximately 87% reduction in token usage under the tested workflow. The result is tied to smart routing, detached processing and a clearly stated workload, not to a universal claim about all AI workloads.

Methodology status: Operational benchmark Internal benchmark, publicly described Tested workflow: financial statement automation Last reviewed: 2026-08-09
approximately 87% reduction in token usageheadline result
~8,500 tokens per statementbaseline per statement
~1,100 tokens per statementoptimized per statement
87%internal benchmark value
Research question

Can smart routing plus detached systems reduce token usage for SME financial statement automation when compared with a heavier AI-first workflow?

Environment / scope

100 SMEs, 12 statements per SME, study date July 2026. The benchmark assumes local inference API fee RM 0/token and separates external API cost from infrastructure amortization.

Methodology
ItemValueWhy it matters
Baseline workflowAI-heavy processing with far more tokens per statement.Represents the reference path.
Optimized workflowSmart routing plus detached systems before LLM escalation.Reduces work sent to the model.
CountedStatement processing tokens, API cost and derived energy estimate.Keeps the benchmark explicit.
Not countedUniversal savings, all hardware variants, all workload types.Avoids overclaiming.
Results
  • Approximately 87% token reduction in the tested workflow.
  • About 6.5M tokens saved across the benchmark model.
  • About RM5,200 in API savings under the stated assumption set.
  • About 47.7 kWh energy and about 85% faster processing in the published study page.
Interpretation

The result supports AINNA's claim that a routed and detached architecture can cut unnecessary model usage. It does not prove the same percentage for other domains, models or infrastructure.

Limitations
  • The benchmark is internal and operational, not peer reviewed.
  • Hardware and pricing assumptions affect the absolute savings.
  • The percentage should not be reused as a universal claim.
  • Further human methodology confirmation is still valuable for external publication.
Related technology

Token Saving Study

Open the underlying benchmark page and assumptions table.

Open page →

Detached Systems

See how deterministic layers reduce model load.

Open page →

Citation information

Suggested citation: AINNA. "NeuralOps Token Efficiency Benchmark." AINNA Research, 2026. Canonical URL: https://ainna.bond/research/neuralops-token-efficiency/

AINNA
CLICK ME

Site Sections

No section data available yet.

Sites with documented sections will appear here.