Can the HY4 vs Kimi K3 benchmark claims pass a source check?
The HY4 vs Kimi K3 coding-benchmark pairing doesn't survive a primary-source check, and coding rank can't settle the best open-source AI for marketing automation. This dated post-mortem logs what verified, what stayed unconfirmed, and how to re-run the check on future vendor claims.
- Platform
- AI coding benchmarks
- Campaign type
- Source check
- Spend range
- No ad spend
- Timeframe
- 2026-08-29
- Primary-source verification
- No verified result
- Verdict
- loss
- Last reviewed
- 2026-08-29
Verdict as of August 29, 2026
The proposed HY4 vs Kimi K3 AI coding benchmark comparison does not pass a primary-source check using the 14-document corpus available for this review. That is a finding about the supplied evidence, not proof that HY4—or a model called Haruka—does not exist.
| Question | Dated finding | Confidence |
|---|---|---|
| Can HY4/Haruka be identified and evaluated from the corpus? | No. The packet provides no confirmed publisher, release, license, model card, evaluation conditions, or benchmark result. | Unconfirmed |
| Can Kimi K3 be confirmed? | Yes, as a named Kimi release. However, the available official passage ends before the full benchmark table and footnotes, leaving no reachable first-party SWE-bench Verified or Aider Polyglot result for K3.[1] | High for model identity; unresolved for the requested benchmark results |
| Does the evidence establish a HY4-versus-K3 coding winner? | No. There are no comparable primary-source results for the two named models. | High |
| Does coding rank identify the best open-source AI for marketing automation? | No. No source in scope tests whether coding-benchmark position predicts performance in marketing-automation workflows. | High within the limits of this review |

The evidence log
The useful output is not a reconstructed leaderboard. It is a record of which claim was checked, where it was checked, and what remained unavailable. “Missing from this packet,” “reported by an aggregator,” and “published by the model vendor” are materially different statuses.
| Claim inspected | Source inspected | What was present | What was missing or limited | Assessment |
|---|---|---|---|---|
| HY4/Haruka is a model that can be compared with Kimi K3 | All 14 documents in the supplied corpus; Onyx AI coding ranking used only as absence-context | The Onyx table names Kimi K2.6. | No HY4 or Haruka entry appeared in the corpus. There was no confirmed identity, publisher, release date, license, model card, or benchmark result. Onyx is not a source for HY4.[2] | Unconfirmed; absence here does not establish nonexistence |
| Kimi K3 is a named release | Official Kimi K3 release note | The release page identifies Kimi K3 and reports product and API information. | The available passage truncates immediately before the full benchmark table and evaluation footnotes.[1] | Confirmed identity; incomplete evaluation record |
| Kimi K3 has a first-party SWE-bench Verified or Aider Polyglot score | Official Kimi K3 release note | A benchmark section is signposted. | The result rows and conditions needed to verify either score are not reachable in the supplied passage.[1] | Unresolved |
| Kimi K3 scored 76.8% on SWE-bench, versus 80.2% for K2.6 | Wan 2.7 benchmark aggregation | Those two figures are displayed by the third-party page. | No evaluation conditions are stated, and the figures cannot be attributed to Moonshot or reconciled against a reachable first-party K3 table.[3] | Low confidence; aggregation only |
| Onyx independently confirms K3’s coding position | Onyx AI 2026 coding ranking | Kimi K2.6 appears in the table. | K3 and HY4 do not appear. A K2.6 row cannot be relabeled as a K3 result.[2] | Unsupported |
| K3 API pricing and cache behavior improve operating economics | Official Kimi K3 release note | The vendor reports $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million output tokens. It also claims a Mooncake cache-hit rate above 90% in coding workloads.[1] | No independent validation appears in the corpus, and coding-workload cache behavior does not establish the same rate for marketing prompts. | Vendor-reported |
| K2.6 evaluation settings can fill the K3 footnote gap | Official Kimi K2.6 release note | K2.6/K2.5 evaluation material specifies temperature 1.0, top-p 1.0, and a 262,144-token context.[4] | Nothing in scope establishes that K3 used the same settings. | Not transferable to K3 |
| SWE-bench Verified offers a timeless, model-independent measure | Epoch AI analysis and discussion containing co-creator commentary | The analysis describes a dataset static since October 2023, with roughly half its issues dating from before 2020, and identifies probable contamination risk. The discussion reports saturation at 93.9% and a subset audit finding test flaws.[5][6] | Contamination is not proven for every model. The audit covered a 27.6% subset rather than the entire benchmark. | Useful benchmark with material interpretation limits |
| K3 architecture, context, license, launch, and arena position are settled facts | WhatLLM secondary tracker | The tracker describes a 2.8-trillion-parameter sparse mixture-of-experts model, a one-million-token context, and a custom Kimi K3 License. It also gives July 16 as the launch date without stating a year and reports a number-one blind frontend-coding arena position at launch.[7] | These details were not corroborated by reachable first-party material in the packet. The custom license also makes “open weight” more precise than automatically calling the model open source. | Medium confidence for architecture, context, and license; low confidence for launch and ranking claims |
The log leaves an asymmetrical comparison. K3 has a confirmed publisher page but an incomplete benchmark record in the material available here. HY4/Haruka lacks even the identity layer required to begin the same check. Combining those gaps into a winner would create evidence rather than evaluate it.
Where the proposed head-to-head breaks
HY4/Haruka never resolves to a verifiable model identity
A model comparison has to begin by resolving the exact object being compared. A name alone is insufficient when it cannot be tied to a publisher, version, dated release, license, model card, or evaluation artifact. None of those identifiers for HY4 or Haruka appears in the supplied corpus.
That warrants the label “not verified here.” It does not warrant “fake,” “nonexistent,” or any similarly categorical conclusion. The model might be documented outside the packet, named differently, privately distributed, or represented by a mistaken shorthand. Until a source resolves that ambiguity, any score attached to the name would be untraceable.
The Onyx ranking does not repair the gap. Its inclusion of Kimi K2.6 provides useful context about what that particular table covers, but its silence on HY4 and K3 is not a benchmark result for either one.[2] A table’s omissions can narrow what the table substantiates; they cannot determine what exists everywhere else.
K3 resolves as a release, then the source chain stops
Kimi K3 clears the first identity check because an official Kimi release note is available. The next checks require the actual benchmark row and its footnotes: benchmark version, subset, harness, provider, sampling settings, context policy, tool configuration, pass criteria, and treatment of failed requests.

In the supplied copy, the official material ends exactly where the full benchmark table and footnotes should begin.[1] That truncation matters. It prevents verification of a first-party SWE-bench Verified result, an Aider Polyglot result, and the conditions under which either result may have been produced.
The Wan 2.7 aggregation supplies numbers—76.8% for K3 and 80.2% for K2.6—but not the missing provenance. Its page does not state the evaluation conditions, and the figures are not attributable to Moonshot within the evidence supplied.[3] They can be recorded as a low-confidence third-party claim. They cannot be promoted to official K3 results.
Nor can settings be borrowed from the neighboring model. The K2.6 release documents temperature 1.0, top-p 1.0, and a 262,144-token context for K2.6/K2.5 evaluations.[4] The shared brand family does not make those K3 settings. Carrying them across would make the comparison look complete while quietly replacing an unknown with an assumption.
Commercial claims belong in a separate column
The official K3 prices are relevant to buyers, but they answer an economics question rather than validating a coding score. The listed rates are $0.30 per million cache-hit input tokens, $3.00 per million cache-miss input tokens, and $15.00 per million output tokens. The same release claims a Mooncake cache-hit rate above 90% in coding workloads.[1]
These are vendor disclosures, not independently verified findings in this corpus. A buyer can use them to estimate a scenario, provided the estimate keeps cache-hit and cache-miss traffic separate. The greater-than-90% figure should not be assumed for campaign generation, audience research, reporting, or other marketing traffic without measurements from those workloads.
A valid coding score still would not select a marketing-automation model
Even a complete and reproducible SWE-bench row would answer a narrower question than the target keyword asks. SWE-bench evaluates software-engineering work derived from repository issues. It does not directly test campaign taxonomy enforcement, ad-policy review, structured media-plan output, connector reliability, approval routing, source-grounded reporting, or recovery from a failed platform request.
Epoch AI’s analysis adds specific cautions. The dataset has been static since October 2023, roughly half its issues predate 2020, and contamination is considered probable, although the analysis does not prove contamination for each individual model.[5] Those limitations do not make the benchmark useless. They make the date, model history, and evaluation method necessary context for interpreting a score.
Saturation and test quality add another boundary. Co-creator commentary in the supplied Hacker News discussion notes a reported 93.9% saturation point and observes that benchmark paradigms eventually saturate. The same thread describes a self-reported audit of a 27.6% subset in which at least 59.4% of the audited problems had flawed tests that rejected functionally correct submissions.[6] Because that audit covers a subset and appears in a discussion thread, it should not be generalized into a defect rate for the full benchmark.
The thread also provides an instructive, anecdotal harness case: a “Pro” variant scored worse than a “Flash” variant partly because it struggled with a custom harness and failed requests were not counted in the comparison.[6] That does not establish how often such distortions occur. It does show why a table row can reflect more than the underlying model: the wrapper, provider behavior, retry policy, timeout handling, and denominator can all affect the recorded result.
| A coding-benchmark row may help establish | It does not establish without additional testing |
|---|---|
| Performance on a named dataset under disclosed evaluation conditions | Accuracy on campaign briefs, media plans, ad-policy checks, or marketing reports |
| Relative results within the same harness and accounting method | Equivalent behavior under a different provider, agent framework, retry policy, or tool stack |
| A reproducible result for a precise model version | Performance for an adjacent version with a similar name |
| Capability on repository-level software tasks | Operational reliability, human-review burden, brand compliance, or cost in a marketing workflow |
No source in scope measures a correlation between coding rank and marketing-automation success. The evidence therefore neither proves nor disproves that a strong coding model could perform well in a particular marketing system. It simply provides no bridge from one result to the other.
What a marketing buyer would still need to test
A task-relevant comparison should use the prompts, tools, data controls, and failure consequences of the intended automation. For example, a buyer evaluating models for media-plan production could create a fixed set of representative briefs and score whether each output preserves supplied budgets, dates, channel constraints, audience exclusions, and required approval fields. That would be a new marketing evaluation, not an interpretation of SWE-bench.
- Record schema-valid outputs separately from outputs that merely look persuasive.
- Count timeouts, refused tool calls, malformed requests, retries, and silent failures in the denominator.
- Test source grounding and whether the model invents campaign IDs, performance figures, policy language, or unsupported recommendations.
- Measure review time and correction burden rather than treating first-pass completion as success.
- Track actual cache behavior, input and output volume, provider charges, and orchestration overhead for the intended workload.
- Review the model license, data-handling terms, deployment options, and restrictions before describing an open-weight release as open source.
Architecture and context-window claims can help determine whether a candidate is worth testing, but they do not replace that test. The secondary WhatLLM tracker describes K3 as a 2.8-trillion-parameter sparse mixture-of-experts model with a one-million-token context and a custom Kimi K3 License.[7] Those are medium-confidence tracker claims in this review. The same page’s July 16 launch date lacks a year, and its reported number-one blind frontend-coding arena position at launch remains low confidence without stronger corroboration.[7]
The available material also cannot support a recommendation among workflow frameworks. A framework shortlist should instead be checked for connector coverage, authentication controls, retry and timeout behavior, observable logs, human approvals, version pinning, provider portability, and the ability to export failed runs for diagnosis.
How to re-run the source check on the next model claim
- Resolve the exact model identity. Capture the publisher, full model name, version, release date, and license. If those cannot be established, label the model unconfirmed rather than guessing.
- Locate the publisher’s primary release. A search snippet, repost, ranking page, or aggregator can point toward a claim but should not silently become the publisher’s evidence.
- Open the benchmark table and footnotes. Confirm that the actual result—not merely a benchmark-section heading—is reachable.
- Capture the evaluation conditions. Note the benchmark version, subset, harness, provider, sampling settings, context policy, tool configuration, pass criteria, retry rules, and handling of failed requests.
- Keep neighboring models separate. Do not transfer settings or scores from K2.6 to K3, from a base model to an agentic variant, or from one provider implementation to another.
- Label provenance in the comparison itself. Distinguish first-party disclosures, independently reproduced results, third-party aggregations, anecdotal reports, and unresolved claims.
- Inspect harness and provider effects. Require failed requests to remain visible and check whether retry or timeout policies alter the denominator.
- Demand evidence for the intended task. A coding result may justify adding a model to a candidate list, but a marketing-automation decision requires a marketing workflow evaluation.
After that procedure, the dated conclusion remains narrow but firm: the HY4-versus-Kimi-K3 pairing does not survive this primary-source check. K3 is identifiable, but its requested first-party benchmark results and conditions are not reachable in the supplied material; HY4/Haruka remains unconfirmed within the corpus. The best open-source AI for marketing automation also remains unanswered because no evidence in scope connects coding rank to success in that work.
References
- Kimi K3 — Kimi
- Best LLMs for Coding 2026 — Onyx AI
- Kimi K3 Benchmarks — Wan 2.7
- Kimi K2.6 — Kimi
- What Skills Does SWE-bench Verified Evaluate? — Epoch AI
- Hacker News thread with SWE-bench co-creator commentary — Hacker News
- Kimi K3 — WhatLLM
Built on this evidence
No Bidding tactic or Creative record currently cites this case file. Compare it against other results in Benchmarks.
Related benchmark reading
Report a corroborating or contradicting result
Seeing something different in your own account? Feed the data-integrity loop instead of leaving an open comment.