Award-winner Ngirinshuti on an AI model that serves Rwanda’s statistical analysis
Award-winner Ngirinshuti on an AI model that serves Rwanda’s statistical analysis
1 October, 2026 •For as long as Rwanda’s Labour Force Survey has existed, someone has had to read a person’s description of their work, written in Kinyarwanda, and match it to a code. Not just any code, but the right line in the International Standard Industrial Classification, one out of thousands of possible entries. Get it wrong, and multiple errors will be made. A mistake can mean that a farmer is recorded as a factory worker, which can affect the accuracy of Rwanda’s national employment statistics and even shift broader trends.
Fidele Ngirinshuti has spent years on the receiving end of that problem. As Labour Force Survey Specialist at the National Institute of Statistics of Rwanda (NISR), he watched experienced analysts lose a full week cross-checking code by hand for each classification system, working through thousands of descriptions collected that quarter, just to be confident the data was right. NISR uses three such systems; ISIC for economic activity, ISCO for occupation, and ISCED for education. Covering all three properly meant at least four analysts, each putting in a week, before a single report could go out.
The seeds of a solution to this challenge emerged when Ngirinshuti participated in a data science training course. The Data Science Capacity Building Initiative is a training programme designed to strengthen Rwanda’s data landscape. Run by the African Institute for Mathematical Sciences (AIMS) in partnership with Cenfri through the Rwanda Economy Digitalisation (RED) Programme, it introduced him to machine learning at a moment when NISR was sitting on nearly a decade of coded survey data going back to 2017. That data, it turned out, was exactly what a model needed to learn from. He said, “When I started the training, I didn’t know what Python was”
What Ngirinshuti and his colleagues built takes economic activity descriptions in Kinyarwanda, English, or French, and predicts the appropriate code, attaching a confidence score to every prediction. High-confidence cases are auto-coded. Lower-confidence cases go to a supervisor. The lowest-confidence cases still go to a human coder. In one test run, 96 percent of records were auto-coded outright correctly, with only 1 percent needing supervisor review and 3 percent needing manual attention.
The same approach was extended beyond economic activity to occupation and education codes, using NISR’s own historical data for each. Work that once needed roughly four analysts a week now runs through one person and a model. NISR’s data cleaning has moved fast enough that reports once due at the end of October can be ready by the end of September, publication schedules permitting.
For now, the tool supports analysts and monitors, cross-checking what enumerators code in the field rather than replacing them. An Android application already exists for eventual use by enumerators themselves, particularly newer staff still learning the classification systems. NISR is approaching this next step carefully, because it wants to be confident in the model before it reaches the field.
Despite this caution the project has already received international recognition! It came through the ITU’s Innovate for Impact award, an initiative that identifies, analyses, and scales practical artificial intelligence solutions for global sustainable development, and Ngirinshuti didn’t expect it. He’s quick to say the model isn’t the most sophisticated one out there, its value lies in how ordinary the ingredients are. Any national statistics office sitting on years of its own coded survey data could, in principle, train something similar.
For him, it is proof the institution built something useful not only for its own reporting timelines, but potentially for statistics offices facing the same coding bottleneck elsewhere. The next test of that will be the Android application ready for field use. Once the model is trained enough to trust with enumerators directly, it could help catch coding errors at the moment data is collected, rather than months later during review.
“If it wasn’t for this data science training, I don’t think I would have had this idea,” he says. He’s candid that the work isn’t finished: “I am still learning how to improve the model. Even though it is working, we still need to enhance it and make sure that we can use it properly.”