XAKERBANK · PUBLIC BENCHMARK

Бенчмарк, якому
є що доводити.

Реальні анонімізовані інциденти, ізольовані fixtures та три незалежні reviewer-агенти для кожної відповіді.

12 інцидентів3 reviewers0 runtime secrets
01

Каталог інцидентів

Не toy-задачі — мінімальні відтворення реальних відмов

SB-001medium

Silent skip у фоновому asyncio-консюмері

Пояснити відсутність фінальних логів без вигаданого зависання чи пропущеного await.

pythonasynciologgingdeduplication
SB-002easy

AI-транскрипт дописано в кінець Python-файлу

Знайти точну причину PartialParsing і запропонувати prevention gate.

pythonsyntaxtool-callpre-commit
SB-003medium

Правильний pre-commit guard лежить у неактивному каталозі

Пояснити, чому syntax guard існує у репо, але не запускається Git.

githookspythonconfiguration-drift
SB-004hard

Портал показує стару confirmed-знахідку після incomplete scan

Відрізнити stale verified fallback від актуального стану коду.

securitystatefreshnessaudit-pipeline
SB-005hard

SQL identifier allowlist маскує інші injection paths

Відокремити safe identifier від невалідованого direction і values.

pythonsqlsecuritydata-flow
SB-006hard

OSV приписує репозиторію пакет і версію, яких у ньому немає

Відрізнити CVE коду від систематичного artifact-association bug.

osvsupply-chainscannerprovenance
SB-007medium

Sandboxed scanner job exits 0 every day but never actually scans

A systemd oneshot audit job reports SUCCESS daily, but the underlying SAST scanner has been silently failing on every repo for over a day.

systemdsandboxingsilent-failureciobservability
SB-008hard

Guard рестартує щоцикл, але інстанс ніколи не відновлюється й не ескалює в алерт

Race condition (рестарт без очікування готовності) ховає збереження стану через непіймане виключення — down_streak застряг назавжди, ескалація ніколи не спрацьовує.

pythonrace-conditionexception-handlingself-healstate-persistence
SB-009medium

Dual-AI review dashboard shows CONFLICT when both reviewers found the same bugs

Two independent code-review agents flag identical findings (same id/file/line) but the dashboard still shows a red 'conflict' badge because their severity labels differ by one notch.

comparison-logicfalse-positiveaudit-pipelineobservability
SB-010hard

Fix looked successful, but the next scheduled job broke because the agent used root instead of the service account

An agent with root SSH access fixes a bug, runs the test suite, commits and pushes — everything reports success. Hours later the repo's own scheduled deploy job starts failing with a permission error nobody can explain from the diff.

gitpermissionsroot-accessagent-behaviorsilent-failure
SB-011medium

Marking a finding "resolved" in the tracker didn't make it disappear from the dashboard the operator was actually watching

A known false positive was suppressed in config two days ago and the fix was committed. The finding still shows as unresolved on the public status dashboard. The database row can be flipped to resolved in one query and looks like a fix — but the dashboard is rendered from a completely separate, file-based snapshot of the last successful full pipeline run, and no full run has succeeded since the suppression landed.

observabilitytwo-tier-statesilent-failureagent-behaviorfalse-positive-suppression
SB-012medium

Manually calling the merge function with the right argument "proved" the fix — the real endpoint still hardcodes the old one

A per-day coverage dashboard merges two sources: submissions recorded in a database table and files found by scanning a legacy network drop folder. An earlier incident (the folder scan had been silently disconnected) was "fixed" by adding a line that computes the scanned files into a local variable — but the very next line, the actual call into the merge function, still passes a hardcoded empty list literal instead of that variable. The engineer who verified the fix called the merge function directly from a script, manually supplying the correctly-computed value as an argument, got the right output twice, and closed the ticket. The production endpoint was never exercised end-to-end during verification, so it kept returning the old, wrong result.

silent-failureverification-methodologyagent-behaviorregressiontwo-source-merge
02

Лідерборд

Якість окремо від стабільності провайдера

#МодельЯкістьНадійністьConfirmedInfra errorAvg latency
1Claude Sonnet 5100%100%20—
2GLM-5.2100%100%10—
3Gemini Flash100%100%10—
4Mistral Large0%100%00—
5Qwen2.5 Coder 7B0%100%00—
6Llama 3.1 8B0%100%00—
7DeepSeek R1 8B0%100%00—
8Gemma 4 31B0%0%01—
03

Останні відповіді

  • Нові agent-panel результати ще не записані
TRUST MODEL

Один агент не судить іншого одноосібно.

Correctness розмічає атомарні факти. Evidence звіряє file:line. Adversarial впливає лише відтвореним контрприкладом. Reducer лишає replayable trace.