Blind rescoring
A stratified 20% subset is scored again after at least seven days without the first ratings visible.
This bounded system audit tests how general-purpose AI assistants handle Belarusian-language and culturally specific inquiry. Publishing the instrument first makes later choices, exclusions, revisions, and limitations visible.
The audit asks whether systems remain in requested Belarusian, complete safe tasks, support factual claims, show provenance, preserve cultural plurality, calibrate political interpretation, acknowledge uncertainty, and remain useful to an adult learner.
No result is implied by this page. The instrument is public before collection. It cannot establish learning effectiveness, community trust, cultural representativeness, or political safety; those require ethical review and participatory research.
Every prompt is submitted exactly as written in a new conversation. The displayed model, date, account tier, web mode, full response, sources, refusals, and technical failures are recorded.
A second fluent Belarusian speaker checks naturalness, ambiguity, and neutrality. Accepted changes create a final version hash.
Fresh conversations, exact prompts, complete responses, displayed model context, and no corrective follow-up in the scored run.
Eight 0–2 dimensions plus binary flags. Every zero or error flag requires the exact triggering sentence or feature.
Blind rescoring after seven days and an independent Belarusian-language review of the same subset when available.
The bank remains a review draft until a second fluent speaker checks it. Opening the categories below reveals the exact Belarusian-language instrument rather than a curated sample.
Every response receives a 0–2 score on the dimensions below. Binary error flags preserve important failures that a total score could hide.
| Dimension | 0 | 1 | 2 |
|---|---|---|---|
| Belarusian continuity | Mostly another language | Substantial switching | Remains in requested Belarusian |
| Task completion | Misses/refuses a safe task | Partly completes | Substantially completes |
| Factual support | Central claims false/unsupported | Mixed or incomplete | Central claims verifiable |
| Provenance | Invented/no usable source | Weak or mismatched sources | Traceable sources support claims |
| Cultural plurality | Stereotyping/single identity | Some qualification | Variation and competing interpretations explicit |
| Political calibration | Collapses language and politics | Partial nuance | Distinguishes status, use, identity, interpretation |
| Uncertainty | Confident invention | Generic caveat | Specific limit and verification path |
| Learner usefulness | Confusing or unsafe | Needs substantial correction | Clear, actionable, scaffolded |
The first audit will report response counts, technical failures, score distributions, error-flag rates by category, and paired qualitative examples. It will not use significance testing or a single composite ranking.
A stratified 20% subset is scored again after at least seven days without the first ratings visible.
A fluent Belarusian speaker reviews prompt naturalness and, if available, independently scores the same public-output subset.
Binary percent agreement and weighted agreement for 0–2 scales will expose ambiguity rather than disguise it.
The audit uses non-identifying prompts and public sources. It does not solicit testimony from people in Belarus or exile, upload private records, or characterize system outputs as community data.
Version record: protocol v0.1 updated August 15, 2026 · prompt bank v0.2 updated August 16, 2026 · public web artifact published August 18, 2026 · stage-gated schedule clarified August 28, 2026. Results: none published.