YOYSTRO

Measurement report · 2026-09-20

Yoystro and six models on the same fifty tasks

Independent measurement on SWE-bench Verified

50 tasks · 18-19 September 2026 · same independent verification

Measured before memory accumulated

Published measurement report of the Yoystro benchmark: same task set, same independent verification, same hardware.

48/50completed tasks (96.0%), the highest across the seven-way comparison
$0.15per completed task, less than half of the nearest model
≈64 mintotal AI run time for all fifty tasks

This report compares Yoystro's full policy against six single model coding assistants that each ran with their own codebase rules file, on the same fifty tasks and before the same independent verification. The measurement was taken before memory accumulated; results are expected to improve as the enterprise memory matures with use, and because that expectation was not measured, no figure for it enters this report.

Yoystro · cognitive AI infrastructure
Yoystro and six models · fifty real tasks · one independent verification
Published edition
SWE-bench Verified
frozen set of 50 tasks

Contents

  1. Summary
  2. 1. Introduction
  3. 2. Method
    2.1 Task set and independent verification · 2.2 Compared parties · 2.3 Fairness principles · 2.4 Cost and time
  4. 3. Results
  5. 4. Advantages
  6. 5. Conclusion
  7. Appendix A. Price classes and cost conversion
  8. Appendix B. Independent verification reports
  9. Appendix C. Produced changes

About this published edition. Under Turkish comparative advertising rules, the tools compared here are identified by class rather than by product or model name (section 2.2). Every figure matches the internal measurement report exactly; nothing was recomputed, rounded or cherry picked.

Nothing is estimated. Where a cell is left empty in this report, the value was not measured; no zero and no estimate is written in its place.

Summary

The same fifty SWE-bench Verified tasks were run by Yoystro and six models on the same hardware, assessed by the same independent verification. The system under measurement is Yoystro, a cognitive AI infrastructure.

This measurement was taken before memory accumulated: the infrastructure ran only once to learn each codebase; results are expected to improve markedly as the enterprise memory matures with use.

Yoystro completed 48 of 50 tasks (96.0%), spent $0.15 per completed task and finished all fifty tasks in roughly 64 minutes of AI run time. The best comparison model reached 46/50 (92.0%) at $0.33 per completed task; the most expensive model completed 42/50 at $1.07 per completed task. The difference does not come from model strength: the model tier Yoystro uses reaches only 43/50 when run on its own, and the same model reaches 48/50 under Yoystro.

Headline results: four measured axes

  • Cost$0.15 per completed task. The cheapest comparison model spent $0.19, the highest scoring one $0.33, the most expensive one $1.07. The same work is completed for four to seven times less money.
  • SpeedFifty tasks in ≈64 minutes of AI run time. The nearest model in the same score band took 78 minutes; the slowest model took 134 minutes to finish six tasks fewer.
  • Work at zero costTwo tasks closed without reaching a model at all, at zero cost.
  • Better work48/50, the highest across the seven-way comparison, and at or tied for the top in all three difficulty classes: easy 22/22, medium 19/20, hard 7/8.

1Introduction

Yoystro is not a coding tool but a cognitive AI infrastructure: an organization's memory of its own codebase, together with an auditable workflow that puts that memory to use.

The word "cognitive" is not a metaphor here; it names the components that were measured: the enterprise memory layer, the rule-based work layer, adaptive routing and verified delivery. Every figure in this report is a joint measurement of those four components; none of them is attributable to the capability of a single model.

The infrastructure addresses three problems. The first is the economics of usage: in continuous work the cost of model calls and the provider side usage allowance become the real constraint on throughput. The second is context continuity: knowledge learned once about a codebase, if it is not carried between sessions, is regenerated and paid for again on every new task. The third is the division of labor: sending work that could be closed in a rule-based way to a model produces both unnecessary cost and unnecessary uncertainty.

The axis of the benchmark follows from those three problems. A completed task rate on its own is not enough. How much the same work costs, how long it takes, and whether knowledge carries across tasks all belong in the same table.

How to read this report

Section 2 establishes the task set, the independent verification, the compared parties and the fairness principles; which rules were applied identically to both sides is stated there. Section 3 gives the measured results. Section 4 sets out which advantage each figure corresponds to. Throughout the report, no estimate has been written into any unmeasured cell.

2Method

2.1 Task set and independent verification

The task set consists of 50 real tasks selected with a fixed seed from the 500 records of SWE-bench Verified, frozen for the duration of the runs, so that figures from different days measure the same tasks.

Independent verification is the SWE-bench harness itself, running inside a sealed container with the same image and the same parser for every party. No result is issued by hand. In the final run all fifty tasks received a result; no task went unmeasured.

Memory state of this measurement

This measurement was taken before memory accumulated: the infrastructure ran only once to learn each codebase; results are expected to improve markedly as the enterprise memory matures with use. Because that expectation was not measured, no figure for it enters this report.

2.2 Compared parties

On one side is Yoystro's full policy. On the other are six single model command line coding assistants, each running with its own codebase rules file at its default settings. In this published edition those six models are named by class. Each compared tool and model configuration is referred to simply as a "model" in this report.

Table 1. The compared parties and their classes
ModelClass
YoystroCognitive AI infrastructure, full policy
Coding assistant ALeading coding assistant, provider 1, top-tier AI model
Coding assistant BLeading coding assistant, provider 1, mid-tier AI model
Coding assistant CLeading coding assistant, provider 1, next-gen AI model
Coding assistant DLeading coding assistant, provider 2, upper-tier AI model
Coding assistant ELeading coding assistant, provider 2, mid-tier AI model
Coding assistant FLeading coding assistant, provider 2, top-tier AI model

All six comparison models ran at the base revision, in a clean working copy, with a fifteen minute ceiling per task.

2.3 Fairness principles

The comparison tools were given no hint that would make the work easier: tests were not applied in advance, the success criterion was not disclosed, and no file path was pointed out. The only input each party received is the publicly available issue text of the task itself. The test names independent verification looks for appear in no party's input, and this is guaranteed at container level.

The same task set, the same independent verification, the same hardware and the same price source apply on both sides.

2.4 Cost and time

Cost. For provider 2 models, token consumption was multiplied by published list prices. For provider 1 models, the total cost reported by the tool itself was used, and that figure was additionally cross checked against list prices. No model was recorded at zero cost: zero would have meant "ran for free".

Time. Time means tool and model working time. Independent verification's container run and any waiting on usage allowance appear in no time column.

3Results

3.1 Main comparison

48/50completed tasks, highest across the seven-way comparison
$0.15per completed task, the lowest
$7.02total cost of all fifty tasks
≈64 mintotal AI run time
Table 2. Yoystro and six models, the same fifty tasks, the same independent verification
RankModelCompletedRateTotal $$/completed$/taskAI run time
1Yoystro48/5096.0%$7.02$0.15$0.14≈64 min
2Coding assistant D46/5092.0%$15.32$0.33$0.3178 min
3Coding assistant A44/5088.0%$33.29$0.76$0.6784 min
4Coding assistant E43/5086.0%$8.22$0.19$0.1672 min
5Coding assistant F42/5084.0%$24.53$0.58$0.4961 min
6Coding assistant B42/5084.0%$24.73$0.59$0.49134 min
7Coding assistant C42/5084.0%$44.92$1.07$0.90122 min
The denominator is 50 for every party. Yoystro's total delivery time, meaning everything from intake to result including verification, is 138 minutes; the AI run time column carries only tool and model working time.
Yoystro 48/50 Coding assistant D 46/50 Coding assistant A 44/50 Coding assistant E 43/50 Coding assistant F 42/50 Coding assistant B 42/50 Coding assistant C 42/50 0 10 20 30 40 50
Figure 1. Completed tasks out of fifty The denominator is 50 for every party. Yoystro leads the seven-way comparison at 48/50; the best of 182 published results on the same set reaches 40/50.
Yoystro $0.15 Coding assistant E $0.19 Coding assistant D $0.33 Coding assistant F $0.58 Coding assistant B $0.59 Coding assistant A $0.76 Coding assistant C $1.07 $0.0 $0.2 $0.4 $0.6 $0.8 $1.0
Figure 2. Cost per completed task (US dollars) Shorter is better. Yoystro is the lowest cost party at $0.15; the highest scoring comparison model spends $0.33 and the most expensive one $1.07.

95% confidence intervals (Wilson method) partially overlap on the score axis (Yoystro 86.5% to 98.9%; nearest model 81.2% to 96.8%). With a sample of fifty tasks this is expected, and the report does not hide it. On the cost axis there is no overlap: Yoystro's $0.15 per completed task is less than half of its nearest competitor.

Published results in context

On the same fifty tasks, the best of 182 results in SWE-bench's own general results archive reaches 40/50. Among the forty seven results that publish cost, the best combination of score and cost completes 40/50 at $0.159 per completed task. Yoystro completes eight more tasks in the same price band.

3.2 Difficulty classes

By the official labels the set splits into 22 easy, 20 medium and 8 hard tasks. The hard class is where a party's real capability shows: on easy tasks all parties converge.

Table 3. Completed tasks by difficulty class
ModelEasy (22)Medium (20)Hard (8)Total
Yoystro22/2219/207/848/50
Coding assistant A21/2219/204/844/50
Coding assistant B21/2217/204/842/50
Coding assistant C21/2217/204/842/50
Coding assistant D22/2216/208/846/50
Coding assistant E19/2217/207/843/50
Coding assistant F21/2217/204/842/50

The hard class splits the field in two: D completes 8/8, Yoystro and E complete 7/8, while A, B, C and F fall to 4/8, that is, to half. In the easy class Yoystro and D both reach the ceiling at 22/22. Yoystro's overall lead does not come from a single class; it comes from standing at or tied for the top in all three, and it produces its two task margin in the medium class.

3.3 Time and context consumption

Table 4. Time and context consumption
ModelTotal AI run timeContext readOutput produced
Yoystro≈64 min10.8 M119 K
Coding assistant D78 min16.5 M130 K
Coding assistant A84 min27.9 M369 K
Coding assistant E72 min16.0 M129 K
Coding assistant F61 min10.4 M63 K
Coding assistant B134 min71.2 M690 K
Coding assistant C122 min19.8 M390 K
Coding assistant F . 42/50 61 min Yoystro . 48/50 ~64 min Coding assistant E . 43/50 72 min Coding assistant D . 46/50 78 min Coding assistant A . 44/50 84 min Coding assistant C . 42/50 122 min Coding assistant B . 42/50 134 min 0 min 35 min 70 min 105 min 140 min
Figure 3. Total AI run time for fifty tasks (minutes) Labels also carry the number of tasks each party completed. One model finishes three minutes sooner but completes six tasks fewer; the slowest model takes more than twice the time to finish six tasks fewer.

One model finishes in less time than Yoystro: F completes 42 tasks in 61 minutes, while Yoystro completes six more tasks in three minutes longer. The slowest model takes more than twice the time to finish six tasks fewer.

The gap on the context side is larger still: Yoystro finishes fifty tasks on 10.8 million tokens of context read, while model B consumes 71.2 million and model A 27.9 million, and both complete six or four tasks fewer. On the output side Yoystro stays at 119 thousand tokens where model B reaches 690 thousand. Not relearning on every task what has been learned once about an organization shows up directly in this column.

3.4 Layers and routing

Two tasks closed at zero cost

The rule-based work layer closed two tasks without sending them to a model at all, at zero cost. In a setup where every task becomes a model call, that outcome cannot be produced by construction. Throughout the run the most expensive model tier was also never called once: the cheaper class finished the work, so the upper step was never needed. This report therefore says nothing about how the most expensive class would behave on this task set.

3.5 Usage allowance efficiency

The number of tasks expected to be completed within a fixed spending window answers the question "how much work gets done on the same budget" directly.

Table 5. Expected completed tasks within a hundred dollar window
Model$/taskRateExpected completed tasks
Yoystro$0.1496.0%≈686
Coding assistant E$0.1686.0%≈538
Coding assistant D$0.3192.0%≈297
Coding assistant F$0.4984.0%≈171
Coding assistant B$0.4984.0%≈171
Coding assistant A$0.6788.0%≈131
Coding assistant C$0.9084.0%≈93
Yoystro ~686 Coding assistant E ~538 Coding assistant D ~297 Coding assistant F ~171 Coding assistant B ~171 Coding assistant A ~131 Coding assistant C ~93 0 175 350 525 700
Figure 4. Expected completed tasks within a hundred dollar window On the same budget Yoystro completes 28% more tasks than its nearest model and roughly 7.3 times more than the most expensive one.

Within the same hundred dollar window Yoystro completes 28% more tasks than its nearest model and roughly 7.3 times more than the most expensive one.

4Advantages

Each of the three headings below rests on a figure measured in section 3. None of them is the achievement of a single model; each is the measured sum of the infrastructure's components.

Components, one outcome

The rule-based work layer closes rule bound work without consulting a model. Adaptive routing finishes every task at the cheapest sufficient level. The enterprise memory layer and verified delivery are parts of the same mechanism; this report names them as components only, without a separate measurement claim. The cost, speed and accuracy advantages all read back to this mechanism.

4.1 Cost: the same work at a quarter to a seventh of the price

At $0.15 per completed task Yoystro is the lowest across the seven-way comparison. The highest scoring comparison model does the same work for $0.33, the most expensive one for $1.07. The total bill for fifty tasks is $7.02 for Yoystro and $44.92 for the most expensive model. This gap is a structure, not a discount: part of the work never reaches a model, and the part that does is closed at the cheapest sufficient level. That the most expensive model tier was never called in this run is the indicator for it.

4.2 Speed: fifty tasks in a little over an hour

All fifty tasks finished in roughly 64 minutes of AI run time. The nearest model in the same score band spent 78 minutes and the third model 84 minutes. The only model that looks faster completes 42 tasks in 61 minutes, while Yoystro completes six more in three minutes longer. The slowest model takes 134 minutes to finish six tasks fewer.

There is a figure on the delivery side as well: total delivery, including queue, preparation and verification, is 138 minutes. Those 138 minutes are therefore not "a change was produced" but "a change was verified".

4.3 Better work: at the top in all three difficulty classes

48/50 is the highest across the seven-way comparison, and the class breakdown shows this does not come from one place: 22/22 in the easy class, 19/20 in the medium class, 7/8 in the hard class. In the hard class one party is one task ahead at 8/8 and one party is level at 7/8; the remaining four fall to 4/8.

The difference does not come from the model's capability, and the indicator sits in the same table: the model tier Yoystro uses reaches only 43/50 when run on its own, and the same model reaches 48/50 under Yoystro. That five task difference was won not by a model upgrade but by how the work is driven.

5Conclusion

On the same fifty tasks Yoystro completes 48/50, spends $0.15 per completed task and finishes the set in roughly 64 minutes of AI run time. The best comparison model completes 46/50 at $0.33 per completed task.

The difference does not come from the model: the model tier Yoystro uses reaches only 43/50 on its own and 48/50 under Yoystro, and two tasks close without reaching a model at all.

In one sentence: on the same fifty tasks Yoystro completes more work, completes it at a quarter to a seventh of the price, completes it in a little over an hour, and does so under the result of independent verification.

Appendix A. Price classes and cost conversion

For provider 2 models, cost is token consumption multiplied by published list prices. Input served from cache is priced separately.

Table 6. List prices used for provider 2 models
Model tierInput ($/M)Cached input ($/M)Output ($/M)
Provider 2, mid class2.000.2012.00
Provider 2, upper class4.000.4020.00
Provider 2, top class10.001.0050.00

For provider 1 models, cost is the total reported by the tool itself, thinking tokens included. That figure was additionally cross checked against list prices; in two models the conversion came out below what the tool reported and in one model above it, and the table used the tool's own figure in all three cases.

Appendix B. Independent verification reports

The table below carries independent verification's result and test results for each of the fifty tasks. FAIL_TO_PASS are the tests a change must make pass; PASS_TO_PASS are the tests it must not break. The task identifiers are the public SWE-bench Verified identifiers. The table covers the main product run only.

48/50tasks independent verification counted as completed
77/78FAIL_TO_PASS tests passed
5970/5970PASS_TO_PASS tests passed, zero regressions
86.4 KBof change in total, 80 files
TaskClassResult FAIL_TO_PASS PASS_TO_PASS Change Files
astropy__astropy-13398hardresolved4/468/6814.7 KB6
astropy__astropy-14508mediumresolved1/1174/1741.5 KB1
astropy__astropy-14539mediumresolved2/246/460.5 KB1
astropy__astropy-14995easyresolved1/1179/1790.6 KB1
astropy__astropy-7166easyresolved1/16/60.7 KB1
django__django-11292mediumresolved1/131/312.8 KB3
django__django-11400hardresolved6/658/583.2 KB3
django__django-11451easyresolved6/645/450.6 KB1
django__django-11532mediumresolved1/1148/1480.4 KB1
django__django-11734mediumunresolved0/1275/2750.6 KB1
django__django-12125easyresolved2/245/450.7 KB1
django__django-12304easyresolved1/117/171.0 KB2
django__django-12325hardresolved2/2201/2011.6 KB2
django__django-13158mediumresolved1/129/290.9 KB1
django__django-13363easyresolved1/176/763.2 KB4
django__django-13401mediumresolved1/132/322.0 KB1
django__django-13406easyresolved3/332/321.4 KB2
django__django-13417easyresolved2/2280/2801.3 KB2
django__django-13551easyresolved2/256/562.0 KB2
django__django-13741easyresolved1/182/823.4 KB3
django__django-14034mediumresolved1/112/121.6 KB1
django__django-14089easyresolved1/143/430.4 KB1
django__django-14915easyresolved1/123/230.4 KB1
django__django-15022mediumresolved3/356/561.1 KB1
django__django-15268hardresolved3/3130/1301.3 KB1
django__django-15503hardresolved2/278/783.2 KB1
django__django-16032mediumresolved2/277/772.2 KB2
django__django-16100easyresolved1/159/591.8 KB1
django__django-16493mediumresolved1/191/910.7 KB1
matplotlib__matplotlib-20859easyresolved1/188/881.0 KB1
matplotlib__matplotlib-25479easyresolved2/2263/2631.1 KB2
mwaskom__seaborn-3069mediumresolved2/294/943.9 KB2
pydata__xarray-3305mediumresolved1/1653/6532.2 KB2
pydata__xarray-3677mediumresolved1/121/211.0 KB2
pydata__xarray-4687mediumresolved1/11717/17172.5 KB3
pylint-dev__pylint-7080mediumresolved1/1120/1200.4 KB1
pytest-dev__pytest-10356hardresolved1/179/792.9 KB2
pytest-dev__pytest-7571mediumresolved1/114/141.3 KB1
scikit-learn__scikit-learn-13124mediumresolved1/160/601.7 KB2
scikit-learn__scikit-learn-13142easyresolved2/254/541.2 KB1
scikit-learn__scikit-learn-14141easyresolved1/12/20.3 KB1
scikit-learn__scikit-learn-25973easyresolved1/172/723.1 KB2
scikit-learn__scikit-learn-26323mediumresolved1/1188/1881.2 KB1
sphinx-doc__sphinx-11510hardno changenot measurednot measured0 KB0
sphinx-doc__sphinx-7454easyresolved1/127/270.7 KB1
sphinx-doc__sphinx-9229hardresolved1/113/132.3 KB2
sphinx-doc__sphinx-9258easyresolved1/145/451.1 KB2
sympy__sympy-14711easyresolved1/12/20.4 KB1
sympy__sympy-16766easyresolved1/17/70.6 KB1
sympy__sympy-23413mediumresolved1/12/21.3 KB1

Appendix C. Produced changes

Every change produced for the fifty tasks is reproduced in a separate document, Verification data: independent verification reports and produced changes. That document carries independent verification's summary result and the full change for each task. The raw files (independent verification reports, run logs and changes) are also distributed as a single verification bundle.