June 30, 2026
Ammar Abdulhamid
To celebrate the historic 25th anniversary of Wikipedia in January 2026, the Hausa Wikimedians User Group organized the Wikipedia @ 25 Writing Contest. The community’s energy was staggering; over 40 participants joined with great enthusiasm. A total of 14,702 articles were created by the contestants on Hausa Wikipedia.
At Hausa Wikimedians UserGroup, a fair and transparent judging process for contests is always a top priority. But with such an incredible volume of contributions, ensuring consistent, high-quality assessment across thousands of entries becomes incredibly difficult to manage manually, so we designed a rigorous, two-part hybrid evaluation system combining automated data analysis with blind human review by a panel of experienced Wikimedia contributors.
A Hybrid Evaluation Approach
The hybrid evaluation process is structured into two distinct phases. Objective Automated Evaluation in the first phase (accounting for 40%) and Human Jury Review in the second (accounting for 60%).
Phase 1: Objective Automated Evaluation
The initial phase focuses on scalable, objective metrics for baseline assessment. In February 2026, we deployed a custom computer program to read and evaluate all the 14,702 submitted articles.
The program graded each article across five structural criteria to calculate raw points:
- Article point: Verifies existence as at time of review, since low-quality articles get deleted fairly quickly.
- Size point: Measures the volume of content added sans references.
-
References point: Measures structural inline citations and footnotes (capped).
-
Sections point: Awards points for sectional and subsectional arrangements (capped).
-
Depth point: Uses heuristics to evaluate content depth and rich details (like media files, graphs, and templates).
Part of Jury Workspace 1 (AR scores)
Part of Jury Workspace 2 (AR scores)
The Fairness Curve: Converting Points to the 40% AR Score
In a heavily popular Wikimedia contest like this one, a simple linear tally poses a challenge: a participant who generates thousands of very brief, often low-quality, articles would always outscore someone who poured immense effort into dozens of deeply researched, highly comprehensive articles.
To ensure structural substance carries proper weight, we institued a "quality buffer" and passed the total raw points through a logarithmic curve (LN) or "diminishing returns" to determine the final 40% Automated Review (AR) score as follows:
The "quality buffer" is quite important here. Writing more articles is certainly appreciated and absolutely increases one's score and keeps them ahead, but it does so at a progressively tapering rate to incentivise slowing down and putting effort into quality too.
After initial review of the raw points, the jury sets the ceiling at 200,000 points (just above the highest participant’s total points of 182,004.5). This was needed since there was no maximum number of articles that a participant can submit for the contest. By using this curve, we ensure that a massive volume of quick, minimal-point stubs cannot completely overwhelm a solid block of high-scoring, well-written articles. It honors sheer output while also appreciating the value of meticulous article writing.
Phase 2: Human Jury Review
While data can measure size and references, it cannot judge the nuances of human language, prose flow, or easily judge contextual quality and coherence. That is where our expert jury comes in. The four-member jury reviewed a representative sample articles.
To determine the sample size (n) of all articles needed for the Jury review, we used a logarithmic scale to ensure a balanced, representative selection for every participant. To maintain grading equity, the evaluation framework also bound the sample size to a minimum of 5 and a maximum of 20 articles per participant (np).
Then for each individual participant, we utilized a Bounded Power-Allocation Stratification to find the number of their articles for the jury review. This scaling ensures that participants who produced exceptionally high volumes of brief articles did not disproportionately exhaust the jury's evaluation capacity at the expense of meticulous writers with lower totals.
We arrived at each participant's allocation (ni) as follows:
- m: The evaluation floor (minimum articles reviewed per person).
- M: The evaluation ceiling (maximum articles reviewed per person).
- Ni: The individual participant's total article count.
- ntotal: The contest Jury bandwith.
- ln(Ni) / ∑ ln(Nj): The logarithmic ratio to normalize and balance the distribution.
Finally, each sampled article was graded by the jury on a qualitative scale from 0 (deleted/nonsensical) to 6 points (excellent).
For increased objectivity, each jury member was assigned an individual evaluation sheet containing a randomized list of articles from multiple authors. After the individual jury evaluations were finalized, the scores were automatically compiled, grouping and summing the qualitative points for each author. This combined tally formed final 60% jury component of the participant's overall score.
Part of scores from Reviewer X and Reviewer Z worksheets shown in above images. Key in column A maps back to author name in coordinator sheet and was only revealed after the reviews were completed. During the review, there was also no row separation between author block of articles.
The contest four-member independent jury panel, which spearheaded the evaluation process, consists of experts Wikimedia editors who bring over 18 years of combined Wikipedia editing experience. The panel was led by Saifullahi Abusufiyan, who also serve as the coordinator of the Wikipedia@25 celebrations and activities in the User Group.
Saifullahi Abusufiyan
Lead Jury Member
Yusuf Sa'adu
Jury Member
Umar Farouq A.
Jury Member
Hassan Abdulhamid
Jury Member
Scaling the Quality (Jury) Score to the 60% Component
With a ceiling of 20 sampled articles and a maximum of 6 points per article, the highest possible raw quality score a participant could achieve from the human jury was 120 points.
To ensure parity with the Automated Review phase, we applied a natural logarithm (LN) to these jury totals. This logarithmic scaling ensures that marginal qualitative differences among the highest-scoring participants do not disproportionately penalize or wipe out editors who achieved strong, highly competitive point totals. The raw jury points were converted into the final 60% Human Review (HR) component as follows:
The final result is the sum of the AR (structural automated metrics) scores and HR (weighed human perspective) scores.
Award Ceremony
The winners will be announced at a special award ceremony which will be organized by Hausa Wikimedians UserGroup in a time to be announced later. For transparency, we are committed to releasing the full result scores after the award ceremony and winners announcement.