← Back to Benchmarks

Can the HY4 vs Kimi K3 benchmark claims pass a source check?

The HY4 vs Kimi K3 coding-benchmark pairing doesn't survive a primary-source check, and coding rank can't settle the best open-source AI for marketing automation. This dated post-mortem logs what verified, what stayed unconfirmed, and how to re-run the check on future vendor claims.

Editorial TeamLOSS
Platform
AI coding benchmarks
Campaign type
Source check
Spend range
No ad spend
Timeframe
2026-08-29
Primary-source verification
No verified result
Verdict
loss
Last reviewed
2026-08-29

Verdict as of August 29, 2026

The proposed HY4 vs Kimi K3 AI coding benchmark comparison does not pass a primary-source check using the 14-document corpus available for this review. That is a finding about the supplied evidence, not proof that HY4—or a model called Haruka—does not exist.

QuestionDated findingConfidence
Can HY4/Haruka be identified and evaluated from the corpus?No. The packet provides no confirmed publisher, release, license, model card, evaluation conditions, or benchmark result.Unconfirmed
Can Kimi K3 be confirmed?Yes, as a named Kimi release. However, the available official passage ends before the full benchmark table and footnotes, leaving no reachable first-party SWE-bench Verified or Aider Polyglot result for K3.[1]High for model identity; unresolved for the requested benchmark results
Does the evidence establish a HY4-versus-K3 coding winner?No. There are no comparable primary-source results for the two named models.High
Does coding rank identify the best open-source AI for marketing automation?No. No source in scope tests whether coding-benchmark position predicts performance in marketing-automation workflows.High within the limits of this review
A benchmark table under a magnifying glass with check and rejection marks beside an evidence card

The evidence log

The useful output is not a reconstructed leaderboard. It is a record of which claim was checked, where it was checked, and what remained unavailable. “Missing from this packet,” “reported by an aggregator,” and “published by the model vendor” are materially different statuses.

Claim inspectedSource inspectedWhat was presentWhat was missing or limitedAssessment
HY4/Haruka is a model that can be compared with Kimi K3All 14 documents in the supplied corpus; Onyx AI coding ranking used only as absence-contextThe Onyx table names Kimi K2.6.No HY4 or Haruka entry appeared in the corpus. There was no confirmed identity, publisher, release date, license, model card, or benchmark result. Onyx is not a source for HY4.[2]Unconfirmed; absence here does not establish nonexistence
Kimi K3 is a named releaseOfficial Kimi K3 release noteThe release page identifies Kimi K3 and reports product and API information.The available passage truncates immediately before the full benchmark table and evaluation footnotes.[1]Confirmed identity; incomplete evaluation record
Kimi K3 has a first-party SWE-bench Verified or Aider Polyglot scoreOfficial Kimi K3 release noteA benchmark section is signposted.The result rows and conditions needed to verify either score are not reachable in the supplied passage.[1]Unresolved
Kimi K3 scored 76.8% on SWE-bench, versus 80.2% for K2.6Wan 2.7 benchmark aggregationThose two figures are displayed by the third-party page.No evaluation conditions are stated, and the figures cannot be attributed to Moonshot or reconciled against a reachable first-party K3 table.[3]Low confidence; aggregation only
Onyx independently confirms K3’s coding positionOnyx AI 2026 coding rankingKimi K2.6 appears in the table.K3 and HY4 do not appear. A K2.6 row cannot be relabeled as a K3 result.[2]Unsupported
K3 API pricing and cache behavior improve operating economicsOfficial Kimi K3 release noteThe vendor reports $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million output tokens. It also claims a Mooncake cache-hit rate above 90% in coding workloads.[1]No independent validation appears in the corpus, and coding-workload cache behavior does not establish the same rate for marketing prompts.Vendor-reported
K2.6 evaluation settings can fill the K3 footnote gapOfficial Kimi K2.6 release noteK2.6/K2.5 evaluation material specifies temperature 1.0, top-p 1.0, and a 262,144-token context.[4]Nothing in scope establishes that K3 used the same settings.Not transferable to K3
SWE-bench Verified offers a timeless, model-independent measureEpoch AI analysis and discussion containing co-creator commentaryThe analysis describes a dataset static since October 2023, with roughly half its issues dating from before 2020, and identifies probable contamination risk. The discussion reports saturation at 93.9% and a subset audit finding test flaws.[5][6]Contamination is not proven for every model. The audit covered a 27.6% subset rather than the entire benchmark.Useful benchmark with material interpretation limits
K3 architecture, context, license, launch, and arena position are settled factsWhatLLM secondary trackerThe tracker describes a 2.8-trillion-parameter sparse mixture-of-experts model, a one-million-token context, and a custom Kimi K3 License. It also gives July 16 as the launch date without stating a year and reports a number-one blind frontend-coding arena position at launch.[7]These details were not corroborated by reachable first-party material in the packet. The custom license also makes “open weight” more precise than automatically calling the model open source.Medium confidence for architecture, context, and license; low confidence for launch and ranking claims

The log leaves an asymmetrical comparison. K3 has a confirmed publisher page but an incomplete benchmark record in the material available here. HY4/Haruka lacks even the identity layer required to begin the same check. Combining those gaps into a winner would create evidence rather than evaluate it.

Where the proposed head-to-head breaks

HY4/Haruka never resolves to a verifiable model identity

A model comparison has to begin by resolving the exact object being compared. A name alone is insufficient when it cannot be tied to a publisher, version, dated release, license, model card, or evaluation artifact. None of those identifiers for HY4 or Haruka appears in the supplied corpus.

That warrants the label “not verified here.” It does not warrant “fake,” “nonexistent,” or any similarly categorical conclusion. The model might be documented outside the packet, named differently, privately distributed, or represented by a mistaken shorthand. Until a source resolves that ambiguity, any score attached to the name would be untraceable.

The Onyx ranking does not repair the gap. Its inclusion of Kimi K2.6 provides useful context about what that particular table covers, but its silence on HY4 and K3 is not a benchmark result for either one.[2] A table’s omissions can narrow what the table substantiates; they cannot determine what exists everywhere else.

K3 resolves as a release, then the source chain stops

Kimi K3 clears the first identity check because an official Kimi release note is available. The next checks require the actual benchmark row and its footnotes: benchmark version, subset, harness, provider, sampling settings, context policy, tool configuration, pass criteria, and treatment of failed requests.

A chain of source documents with a broken link and a final document cut off at the bottom

In the supplied copy, the official material ends exactly where the full benchmark table and footnotes should begin.[1] That truncation matters. It prevents verification of a first-party SWE-bench Verified result, an Aider Polyglot result, and the conditions under which either result may have been produced.

The Wan 2.7 aggregation supplies numbers—76.8% for K3 and 80.2% for K2.6—but not the missing provenance. Its page does not state the evaluation conditions, and the figures are not attributable to Moonshot within the evidence supplied.[3] They can be recorded as a low-confidence third-party claim. They cannot be promoted to official K3 results.

Nor can settings be borrowed from the neighboring model. The K2.6 release documents temperature 1.0, top-p 1.0, and a 262,144-token context for K2.6/K2.5 evaluations.[4] The shared brand family does not make those K3 settings. Carrying them across would make the comparison look complete while quietly replacing an unknown with an assumption.

Commercial claims belong in a separate column

The official K3 prices are relevant to buyers, but they answer an economics question rather than validating a coding score. The listed rates are $0.30 per million cache-hit input tokens, $3.00 per million cache-miss input tokens, and $15.00 per million output tokens. The same release claims a Mooncake cache-hit rate above 90% in coding workloads.[1]

These are vendor disclosures, not independently verified findings in this corpus. A buyer can use them to estimate a scenario, provided the estimate keeps cache-hit and cache-miss traffic separate. The greater-than-90% figure should not be assumed for campaign generation, audience research, reporting, or other marketing traffic without measurements from those workloads.

A valid coding score still would not select a marketing-automation model

Even a complete and reproducible SWE-bench row would answer a narrower question than the target keyword asks. SWE-bench evaluates software-engineering work derived from repository issues. It does not directly test campaign taxonomy enforcement, ad-policy review, structured media-plan output, connector reliability, approval routing, source-grounded reporting, or recovery from a failed platform request.

Epoch AI’s analysis adds specific cautions. The dataset has been static since October 2023, roughly half its issues predate 2020, and contamination is considered probable, although the analysis does not prove contamination for each individual model.[5] Those limitations do not make the benchmark useless. They make the date, model history, and evaluation method necessary context for interpreting a score.

Saturation and test quality add another boundary. Co-creator commentary in the supplied Hacker News discussion notes a reported 93.9% saturation point and observes that benchmark paradigms eventually saturate. The same thread describes a self-reported audit of a 27.6% subset in which at least 59.4% of the audited problems had flawed tests that rejected functionally correct submissions.[6] Because that audit covers a subset and appears in a discussion thread, it should not be generalized into a defect rate for the full benchmark.

The thread also provides an instructive, anecdotal harness case: a “Pro” variant scored worse than a “Flash” variant partly because it struggled with a custom harness and failed requests were not counted in the comparison.[6] That does not establish how often such distortions occur. It does show why a table row can reflect more than the underlying model: the wrapper, provider behavior, retry policy, timeout handling, and denominator can all affect the recorded result.

A coding-benchmark row may help establishIt does not establish without additional testing
Performance on a named dataset under disclosed evaluation conditionsAccuracy on campaign briefs, media plans, ad-policy checks, or marketing reports
Relative results within the same harness and accounting methodEquivalent behavior under a different provider, agent framework, retry policy, or tool stack
A reproducible result for a precise model versionPerformance for an adjacent version with a similar name
Capability on repository-level software tasksOperational reliability, human-review burden, brand compliance, or cost in a marketing workflow

No source in scope measures a correlation between coding rank and marketing-automation success. The evidence therefore neither proves nor disproves that a strong coding model could perform well in a particular marketing system. It simply provides no bridge from one result to the other.

What a marketing buyer would still need to test

A task-relevant comparison should use the prompts, tools, data controls, and failure consequences of the intended automation. For example, a buyer evaluating models for media-plan production could create a fixed set of representative briefs and score whether each output preserves supplied budgets, dates, channel constraints, audience exclusions, and required approval fields. That would be a new marketing evaluation, not an interpretation of SWE-bench.

  • Record schema-valid outputs separately from outputs that merely look persuasive.
  • Count timeouts, refused tool calls, malformed requests, retries, and silent failures in the denominator.
  • Test source grounding and whether the model invents campaign IDs, performance figures, policy language, or unsupported recommendations.
  • Measure review time and correction burden rather than treating first-pass completion as success.
  • Track actual cache behavior, input and output volume, provider charges, and orchestration overhead for the intended workload.
  • Review the model license, data-handling terms, deployment options, and restrictions before describing an open-weight release as open source.

Architecture and context-window claims can help determine whether a candidate is worth testing, but they do not replace that test. The secondary WhatLLM tracker describes K3 as a 2.8-trillion-parameter sparse mixture-of-experts model with a one-million-token context and a custom Kimi K3 License.[7] Those are medium-confidence tracker claims in this review. The same page’s July 16 launch date lacks a year, and its reported number-one blind frontend-coding arena position at launch remains low confidence without stronger corroboration.[7]

The available material also cannot support a recommendation among workflow frameworks. A framework shortlist should instead be checked for connector coverage, authentication controls, retry and timeout behavior, observable logs, human approvals, version pinning, provider portability, and the ability to export failed runs for diagnosis.

How to re-run the source check on the next model claim

  1. Resolve the exact model identity. Capture the publisher, full model name, version, release date, and license. If those cannot be established, label the model unconfirmed rather than guessing.
  2. Locate the publisher’s primary release. A search snippet, repost, ranking page, or aggregator can point toward a claim but should not silently become the publisher’s evidence.
  3. Open the benchmark table and footnotes. Confirm that the actual result—not merely a benchmark-section heading—is reachable.
  4. Capture the evaluation conditions. Note the benchmark version, subset, harness, provider, sampling settings, context policy, tool configuration, pass criteria, retry rules, and handling of failed requests.
  5. Keep neighboring models separate. Do not transfer settings or scores from K2.6 to K3, from a base model to an agentic variant, or from one provider implementation to another.
  6. Label provenance in the comparison itself. Distinguish first-party disclosures, independently reproduced results, third-party aggregations, anecdotal reports, and unresolved claims.
  7. Inspect harness and provider effects. Require failed requests to remain visible and check whether retry or timeout policies alter the denominator.
  8. Demand evidence for the intended task. A coding result may justify adding a model to a candidate list, but a marketing-automation decision requires a marketing workflow evaluation.

After that procedure, the dated conclusion remains narrow but firm: the HY4-versus-Kimi-K3 pairing does not survive this primary-source check. K3 is identifiable, but its requested first-party benchmark results and conditions are not reachable in the supplied material; HY4/Haruka remains unconfirmed within the corpus. The best open-source AI for marketing automation also remains unanswered because no evidence in scope connects coding rank to success in that work.

References

  1. Kimi K3 — Kimi
  2. Best LLMs for Coding 2026 — Onyx AI
  3. Kimi K3 Benchmarks — Wan 2.7
  4. Kimi K2.6 — Kimi
  5. What Skills Does SWE-bench Verified Evaluate? — Epoch AI
  6. Hacker News thread with SWE-bench co-creator commentary — Hacker News
  7. Kimi K3 — WhatLLM

No Bidding tactic or Creative record currently cites this case file. Compare it against other results in Benchmarks.

Related benchmark reading

Report a corroborating or contradicting result

Seeing something different in your own account? Feed the data-integrity loop instead of leaving an open comment.