Summary
The same fifty SWE-bench Verified tasks were run by Yoystro and six models on the same hardware, assessed by the same independent verification. The system under measurement is Yoystro, a cognitive AI infrastructure.
This measurement was taken before memory accumulated: the infrastructure ran only once to learn each codebase; results are expected to improve markedly as the enterprise memory matures with use.
Yoystro completed 48 of 50 tasks (96.0%), spent $0.15 per completed task and finished all fifty tasks in roughly 64 minutes of AI run time. The best comparison model reached 46/50 (92.0%) at $0.33 per completed task; the most expensive model completed 42/50 at $1.07 per completed task. The difference does not come from model strength: the model tier Yoystro uses reaches only 43/50 when run on its own, and the same model reaches 48/50 under Yoystro.
Headline results: four measured axes
- Cost$0.15 per completed task. The cheapest comparison model spent $0.19, the highest scoring one $0.33, the most expensive one $1.07. The same work is completed for four to seven times less money.
- SpeedFifty tasks in ≈64 minutes of AI run time. The nearest model in the same score band took 78 minutes; the slowest model took 134 minutes to finish six tasks fewer.
- Work at zero costTwo tasks closed without reaching a model at all, at zero cost.
- Better work48/50, the highest across the seven-way comparison, and at or tied for the top in all three difficulty classes: easy 22/22, medium 19/20, hard 7/8.
1Introduction
Yoystro is not a coding tool but a cognitive AI infrastructure: an organization's memory of its own codebase, together with an auditable workflow that puts that memory to use.
The word "cognitive" is not a metaphor here; it names the components that were measured: the enterprise memory layer, the rule-based work layer, adaptive routing and verified delivery. Every figure in this report is a joint measurement of those four components; none of them is attributable to the capability of a single model.
The infrastructure addresses three problems. The first is the economics of usage: in continuous work the cost of model calls and the provider side usage allowance become the real constraint on throughput. The second is context continuity: knowledge learned once about a codebase, if it is not carried between sessions, is regenerated and paid for again on every new task. The third is the division of labor: sending work that could be closed in a rule-based way to a model produces both unnecessary cost and unnecessary uncertainty.
The axis of the benchmark follows from those three problems. A completed task rate on its own is not enough. How much the same work costs, how long it takes, and whether knowledge carries across tasks all belong in the same table.
How to read this report
Section 2 establishes the task set, the independent verification, the compared parties and the fairness principles; which rules were applied identically to both sides is stated there. Section 3 gives the measured results. Section 4 sets out which advantage each figure corresponds to. Throughout the report, no estimate has been written into any unmeasured cell.
2Method
2.1 Task set and independent verification
The task set consists of 50 real tasks selected with a fixed seed from the 500 records of SWE-bench Verified, frozen for the duration of the runs, so that figures from different days measure the same tasks.
Independent verification is the SWE-bench harness itself, running inside a sealed container with the same image and the same parser for every party. No result is issued by hand. In the final run all fifty tasks received a result; no task went unmeasured.
Memory state of this measurement
This measurement was taken before memory accumulated: the infrastructure ran only once to learn each codebase; results are expected to improve markedly as the enterprise memory matures with use. Because that expectation was not measured, no figure for it enters this report.
2.2 Compared parties
On one side is Yoystro's full policy. On the other are six single model command line coding assistants, each running with its own codebase rules file at its default settings. In this published edition those six models are named by class. Each compared tool and model configuration is referred to simply as a "model" in this report.
| Model | Class |
|---|---|
| Yoystro | Cognitive AI infrastructure, full policy |
| Coding assistant A | Leading coding assistant, provider 1, top-tier AI model |
| Coding assistant B | Leading coding assistant, provider 1, mid-tier AI model |
| Coding assistant C | Leading coding assistant, provider 1, next-gen AI model |
| Coding assistant D | Leading coding assistant, provider 2, upper-tier AI model |
| Coding assistant E | Leading coding assistant, provider 2, mid-tier AI model |
| Coding assistant F | Leading coding assistant, provider 2, top-tier AI model |
All six comparison models ran at the base revision, in a clean working copy, with a fifteen minute ceiling per task.
2.3 Fairness principles
The comparison tools were given no hint that would make the work easier: tests were not applied in advance, the success criterion was not disclosed, and no file path was pointed out. The only input each party received is the publicly available issue text of the task itself. The test names independent verification looks for appear in no party's input, and this is guaranteed at container level.
The same task set, the same independent verification, the same hardware and the same price source apply on both sides.
2.4 Cost and time
Cost. For provider 2 models, token consumption was multiplied by published list prices. For provider 1 models, the total cost reported by the tool itself was used, and that figure was additionally cross checked against list prices. No model was recorded at zero cost: zero would have meant "ran for free".
Time. Time means tool and model working time. Independent verification's container run and any waiting on usage allowance appear in no time column.
3Results
3.1 Main comparison
| Rank | Model | Completed | Rate | Total $ | $/completed | $/task | AI run time |
|---|---|---|---|---|---|---|---|
| 1 | Yoystro | 48/50 | 96.0% | $7.02 | $0.15 | $0.14 | ≈64 min |
| 2 | Coding assistant D | 46/50 | 92.0% | $15.32 | $0.33 | $0.31 | 78 min |
| 3 | Coding assistant A | 44/50 | 88.0% | $33.29 | $0.76 | $0.67 | 84 min |
| 4 | Coding assistant E | 43/50 | 86.0% | $8.22 | $0.19 | $0.16 | 72 min |
| 5 | Coding assistant F | 42/50 | 84.0% | $24.53 | $0.58 | $0.49 | 61 min |
| 6 | Coding assistant B | 42/50 | 84.0% | $24.73 | $0.59 | $0.49 | 134 min |
| 7 | Coding assistant C | 42/50 | 84.0% | $44.92 | $1.07 | $0.90 | 122 min |
| The denominator is 50 for every party. Yoystro's total delivery time, meaning everything from intake to result including verification, is 138 minutes; the AI run time column carries only tool and model working time. | |||||||
95% confidence intervals (Wilson method) partially overlap on the score axis (Yoystro 86.5% to 98.9%; nearest model 81.2% to 96.8%). With a sample of fifty tasks this is expected, and the report does not hide it. On the cost axis there is no overlap: Yoystro's $0.15 per completed task is less than half of its nearest competitor.
Published results in context
On the same fifty tasks, the best of 182 results in SWE-bench's own general results archive reaches 40/50. Among the forty seven results that publish cost, the best combination of score and cost completes 40/50 at $0.159 per completed task. Yoystro completes eight more tasks in the same price band.
3.2 Difficulty classes
By the official labels the set splits into 22 easy, 20 medium and 8 hard tasks. The hard class is where a party's real capability shows: on easy tasks all parties converge.
| Model | Easy (22) | Medium (20) | Hard (8) | Total |
|---|---|---|---|---|
| Yoystro | 22/22 | 19/20 | 7/8 | 48/50 |
| Coding assistant A | 21/22 | 19/20 | 4/8 | 44/50 |
| Coding assistant B | 21/22 | 17/20 | 4/8 | 42/50 |
| Coding assistant C | 21/22 | 17/20 | 4/8 | 42/50 |
| Coding assistant D | 22/22 | 16/20 | 8/8 | 46/50 |
| Coding assistant E | 19/22 | 17/20 | 7/8 | 43/50 |
| Coding assistant F | 21/22 | 17/20 | 4/8 | 42/50 |
The hard class splits the field in two: D completes 8/8, Yoystro and E complete 7/8, while A, B, C and F fall to 4/8, that is, to half. In the easy class Yoystro and D both reach the ceiling at 22/22. Yoystro's overall lead does not come from a single class; it comes from standing at or tied for the top in all three, and it produces its two task margin in the medium class.
3.3 Time and context consumption
| Model | Total AI run time | Context read | Output produced |
|---|---|---|---|
| Yoystro | ≈64 min | 10.8 M | 119 K |
| Coding assistant D | 78 min | 16.5 M | 130 K |
| Coding assistant A | 84 min | 27.9 M | 369 K |
| Coding assistant E | 72 min | 16.0 M | 129 K |
| Coding assistant F | 61 min | 10.4 M | 63 K |
| Coding assistant B | 134 min | 71.2 M | 690 K |
| Coding assistant C | 122 min | 19.8 M | 390 K |
One model finishes in less time than Yoystro: F completes 42 tasks in 61 minutes, while Yoystro completes six more tasks in three minutes longer. The slowest model takes more than twice the time to finish six tasks fewer.
The gap on the context side is larger still: Yoystro finishes fifty tasks on 10.8 million tokens of context read, while model B consumes 71.2 million and model A 27.9 million, and both complete six or four tasks fewer. On the output side Yoystro stays at 119 thousand tokens where model B reaches 690 thousand. Not relearning on every task what has been learned once about an organization shows up directly in this column.
3.4 Layers and routing
Two tasks closed at zero cost
The rule-based work layer closed two tasks without sending them to a model at all, at zero cost. In a setup where every task becomes a model call, that outcome cannot be produced by construction. Throughout the run the most expensive model tier was also never called once: the cheaper class finished the work, so the upper step was never needed. This report therefore says nothing about how the most expensive class would behave on this task set.
3.5 Usage allowance efficiency
The number of tasks expected to be completed within a fixed spending window answers the question "how much work gets done on the same budget" directly.
| Model | $/task | Rate | Expected completed tasks |
|---|---|---|---|
| Yoystro | $0.14 | 96.0% | ≈686 |
| Coding assistant E | $0.16 | 86.0% | ≈538 |
| Coding assistant D | $0.31 | 92.0% | ≈297 |
| Coding assistant F | $0.49 | 84.0% | ≈171 |
| Coding assistant B | $0.49 | 84.0% | ≈171 |
| Coding assistant A | $0.67 | 88.0% | ≈131 |
| Coding assistant C | $0.90 | 84.0% | ≈93 |
Within the same hundred dollar window Yoystro completes 28% more tasks than its nearest model and roughly 7.3 times more than the most expensive one.
4Advantages
Each of the three headings below rests on a figure measured in section 3. None of them is the achievement of a single model; each is the measured sum of the infrastructure's components.
Components, one outcome
The rule-based work layer closes rule bound work without consulting a model. Adaptive routing finishes every task at the cheapest sufficient level. The enterprise memory layer and verified delivery are parts of the same mechanism; this report names them as components only, without a separate measurement claim. The cost, speed and accuracy advantages all read back to this mechanism.
4.1 Cost: the same work at a quarter to a seventh of the price
At $0.15 per completed task Yoystro is the lowest across the seven-way comparison. The highest scoring comparison model does the same work for $0.33, the most expensive one for $1.07. The total bill for fifty tasks is $7.02 for Yoystro and $44.92 for the most expensive model. This gap is a structure, not a discount: part of the work never reaches a model, and the part that does is closed at the cheapest sufficient level. That the most expensive model tier was never called in this run is the indicator for it.
4.2 Speed: fifty tasks in a little over an hour
All fifty tasks finished in roughly 64 minutes of AI run time. The nearest model in the same score band spent 78 minutes and the third model 84 minutes. The only model that looks faster completes 42 tasks in 61 minutes, while Yoystro completes six more in three minutes longer. The slowest model takes 134 minutes to finish six tasks fewer.
There is a figure on the delivery side as well: total delivery, including queue, preparation and verification, is 138 minutes. Those 138 minutes are therefore not "a change was produced" but "a change was verified".
4.3 Better work: at the top in all three difficulty classes
48/50 is the highest across the seven-way comparison, and the class breakdown shows this does not come from one place: 22/22 in the easy class, 19/20 in the medium class, 7/8 in the hard class. In the hard class one party is one task ahead at 8/8 and one party is level at 7/8; the remaining four fall to 4/8.
The difference does not come from the model's capability, and the indicator sits in the same table: the model tier Yoystro uses reaches only 43/50 when run on its own, and the same model reaches 48/50 under Yoystro. That five task difference was won not by a model upgrade but by how the work is driven.
5Conclusion
On the same fifty tasks Yoystro completes 48/50, spends $0.15 per completed task and finishes the set in roughly 64 minutes of AI run time. The best comparison model completes 46/50 at $0.33 per completed task.
The difference does not come from the model: the model tier Yoystro uses reaches only 43/50 on its own and 48/50 under Yoystro, and two tasks close without reaching a model at all.
In one sentence: on the same fifty tasks Yoystro completes more work, completes it at a quarter to a seventh of the price, completes it in a little over an hour, and does so under the result of independent verification.
Appendix A. Price classes and cost conversion
For provider 2 models, cost is token consumption multiplied by published list prices. Input served from cache is priced separately.
| Model tier | Input ($/M) | Cached input ($/M) | Output ($/M) |
|---|---|---|---|
| Provider 2, mid class | 2.00 | 0.20 | 12.00 |
| Provider 2, upper class | 4.00 | 0.40 | 20.00 |
| Provider 2, top class | 10.00 | 1.00 | 50.00 |
For provider 1 models, cost is the total reported by the tool itself, thinking tokens included. That figure was additionally cross checked against list prices; in two models the conversion came out below what the tool reported and in one model above it, and the table used the tool's own figure in all three cases.
Appendix B. Independent verification reports
The table below carries independent verification's result and test results for each of the fifty tasks. FAIL_TO_PASS are the tests a change must make pass; PASS_TO_PASS are the tests it must not break. The task identifiers are the public SWE-bench Verified identifiers. The table covers the main product run only.
| Task | Class | Result | FAIL_TO_PASS | PASS_TO_PASS | Change | Files |
|---|---|---|---|---|---|---|
| astropy__astropy-13398 | hard | resolved | 4/4 | 68/68 | 14.7 KB | 6 |
| astropy__astropy-14508 | medium | resolved | 1/1 | 174/174 | 1.5 KB | 1 |
| astropy__astropy-14539 | medium | resolved | 2/2 | 46/46 | 0.5 KB | 1 |
| astropy__astropy-14995 | easy | resolved | 1/1 | 179/179 | 0.6 KB | 1 |
| astropy__astropy-7166 | easy | resolved | 1/1 | 6/6 | 0.7 KB | 1 |
| django__django-11292 | medium | resolved | 1/1 | 31/31 | 2.8 KB | 3 |
| django__django-11400 | hard | resolved | 6/6 | 58/58 | 3.2 KB | 3 |
| django__django-11451 | easy | resolved | 6/6 | 45/45 | 0.6 KB | 1 |
| django__django-11532 | medium | resolved | 1/1 | 148/148 | 0.4 KB | 1 |
| django__django-11734 | medium | unresolved | 0/1 | 275/275 | 0.6 KB | 1 |
| django__django-12125 | easy | resolved | 2/2 | 45/45 | 0.7 KB | 1 |
| django__django-12304 | easy | resolved | 1/1 | 17/17 | 1.0 KB | 2 |
| django__django-12325 | hard | resolved | 2/2 | 201/201 | 1.6 KB | 2 |
| django__django-13158 | medium | resolved | 1/1 | 29/29 | 0.9 KB | 1 |
| django__django-13363 | easy | resolved | 1/1 | 76/76 | 3.2 KB | 4 |
| django__django-13401 | medium | resolved | 1/1 | 32/32 | 2.0 KB | 1 |
| django__django-13406 | easy | resolved | 3/3 | 32/32 | 1.4 KB | 2 |
| django__django-13417 | easy | resolved | 2/2 | 280/280 | 1.3 KB | 2 |
| django__django-13551 | easy | resolved | 2/2 | 56/56 | 2.0 KB | 2 |
| django__django-13741 | easy | resolved | 1/1 | 82/82 | 3.4 KB | 3 |
| django__django-14034 | medium | resolved | 1/1 | 12/12 | 1.6 KB | 1 |
| django__django-14089 | easy | resolved | 1/1 | 43/43 | 0.4 KB | 1 |
| django__django-14915 | easy | resolved | 1/1 | 23/23 | 0.4 KB | 1 |
| django__django-15022 | medium | resolved | 3/3 | 56/56 | 1.1 KB | 1 |
| django__django-15268 | hard | resolved | 3/3 | 130/130 | 1.3 KB | 1 |
| django__django-15503 | hard | resolved | 2/2 | 78/78 | 3.2 KB | 1 |
| django__django-16032 | medium | resolved | 2/2 | 77/77 | 2.2 KB | 2 |
| django__django-16100 | easy | resolved | 1/1 | 59/59 | 1.8 KB | 1 |
| django__django-16493 | medium | resolved | 1/1 | 91/91 | 0.7 KB | 1 |
| matplotlib__matplotlib-20859 | easy | resolved | 1/1 | 88/88 | 1.0 KB | 1 |
| matplotlib__matplotlib-25479 | easy | resolved | 2/2 | 263/263 | 1.1 KB | 2 |
| mwaskom__seaborn-3069 | medium | resolved | 2/2 | 94/94 | 3.9 KB | 2 |
| pydata__xarray-3305 | medium | resolved | 1/1 | 653/653 | 2.2 KB | 2 |
| pydata__xarray-3677 | medium | resolved | 1/1 | 21/21 | 1.0 KB | 2 |
| pydata__xarray-4687 | medium | resolved | 1/1 | 1717/1717 | 2.5 KB | 3 |
| pylint-dev__pylint-7080 | medium | resolved | 1/1 | 120/120 | 0.4 KB | 1 |
| pytest-dev__pytest-10356 | hard | resolved | 1/1 | 79/79 | 2.9 KB | 2 |
| pytest-dev__pytest-7571 | medium | resolved | 1/1 | 14/14 | 1.3 KB | 1 |
| scikit-learn__scikit-learn-13124 | medium | resolved | 1/1 | 60/60 | 1.7 KB | 2 |
| scikit-learn__scikit-learn-13142 | easy | resolved | 2/2 | 54/54 | 1.2 KB | 1 |
| scikit-learn__scikit-learn-14141 | easy | resolved | 1/1 | 2/2 | 0.3 KB | 1 |
| scikit-learn__scikit-learn-25973 | easy | resolved | 1/1 | 72/72 | 3.1 KB | 2 |
| scikit-learn__scikit-learn-26323 | medium | resolved | 1/1 | 188/188 | 1.2 KB | 1 |
| sphinx-doc__sphinx-11510 | hard | no change | not measured | not measured | 0 KB | 0 |
| sphinx-doc__sphinx-7454 | easy | resolved | 1/1 | 27/27 | 0.7 KB | 1 |
| sphinx-doc__sphinx-9229 | hard | resolved | 1/1 | 13/13 | 2.3 KB | 2 |
| sphinx-doc__sphinx-9258 | easy | resolved | 1/1 | 45/45 | 1.1 KB | 2 |
| sympy__sympy-14711 | easy | resolved | 1/1 | 2/2 | 0.4 KB | 1 |
| sympy__sympy-16766 | easy | resolved | 1/1 | 7/7 | 0.6 KB | 1 |
| sympy__sympy-23413 | medium | resolved | 1/1 | 2/2 | 1.3 KB | 1 |
Appendix C. Produced changes
Every change produced for the fifty tasks is reproduced in a separate document, Verification data: independent verification reports and produced changes. That document carries independent verification's summary result and the full change for each task. The raw files (independent verification reports, run logs and changes) are also distributed as a single verification bundle.