Thank you for visiting the HLA website, we are updating the site to reflect changes in our organisation. Contact us if you have any questions. Contact us

Humanitarian AI research in action: a humanitarian technologist’s view from Colombia  

How can academic AI research directly shape and inform humanitarian response? 

In 2025, the Humanitarian Leadership Academy (HLA) and Data Friendly Space (DFS) led the world’s first global study into how humanitarians are using AI, forming a global baseline that is informing and shaping policy, governance, research and sector-wide discourse.

In this interview, we connect with Joshua Beretta from Kobo, who has just successfully completed his master’s research in humanitarian AI in the context of Colombia, which builds on some of the themes that emerged in the HLA/DFS study. 

The HLA’s Ka Man Parkinson, humanitarian AI research co-lead, caught up with Joshua -12 months after they first connected – to hear more about his research journey and work.  


Introducing Joshua and his work 

I’m the Humanitarian Programmes Lead at Kobo. I started on the team about six years ago as a backend developer, and I’ve taken on many different roles internally, but have recently been building up our humanitarian projects, and now our humanitarian team and emergency response capabilities.  

I do a lot of work internally on our innovation projects, and these days I focus on how to use AI responsibly in humanitarian contexts specifically – but also more generally in the nexus space of humanitarian, development and peacebuilding work, and how technology nonprofits like Kobo can incorporate and use AI responsibly when it’s serving millions of people. 

I have a mechanical engineering background – and AI was starting to become a much bigger deal in the public’s eye around the time I graduated, 2013/2014, with a lot of the developments in image processing and computer vision, and the work of DeepMind. I always had it as a side interest, and was always following the developments of it.  

Then my career started moving more in the nonprofit technology direction, and I started seeing, especially since the launch of ChatGPT, the massive potential AI has to bring its capabilities into the demanding, resource-scarce environment of humanitarian response – but also that this now brings significant risks to the sector, and to the people the sector aims to serve. 

Three people in vests sit and kneel around a table with laptops, engaged in discussion at an indoor event. The table is covered with a green and white floral-patterned cloth.
Pictured: Joshua Beretta (right) facilitating a KoboToolbox training and co-design workshop in Bogotá for local organisations. Image credit: Kobo 

 Why Colombia 

Joshua’s master’s thesis split the humanitarian AI question into two groups: practitioners working directly with affected populations, and the technical experts building the tools they use. He chose to research it somewhere the conversation rarely reaches. 

“I wanted to see not just a Global North perspective, but specifically the Global South and a non-English environment – the more complex dynamics of AI in low-resource language environments, in contexts where there’s very little information the AIs have been trained on.  

I spent six weeks based in Bogotá. I did as much of the practitioner interviews in person as I could, and in Spanish. My Spanish isn’t perfect, but I wanted people to be able to express themselves in the language they felt most comfortable in.”

The linguistic diversity of Colombia and the data challenge 

It is one thing to know, in the abstract, that generative AI are language models. It can be another to sit with what that means in practice – how directly a language’s presence online decides whose needs an AI system can serve, and whose it can’t. 

“Colombia is linguistically diverse – not just Colombia’s own Indigenous languages, but the migration coming in from Venezuela bringing even more linguistic diversity into an already under-served landscape.  

Spanish has very good support within AI models. But almost every other minority or Indigenous language has almost no support – or the support it has is so minimal that it struggles to understand local dialects, or the switching between a local language and Spanish, or English, or Portuguese. 

Of the roughly sixty Indigenous languages spoken across the two countries, none are covered by Meta’s older NLLB-200 translation model, and only two – Arhuaco and Wayuu – appear in the evaluation set for its newer Omnilingual MT model, despite that model’s headline claim of supporting sixteen hundred languages.

Speech recognition fares a little better, with 12 of Colombia’s 29 and 16 of Venezuela’s 43 Indigenous languages picked up by Meta’s Omnilingual ASR model – but most of the training text that does exist for these languages is Bible translation, not everyday language, therefore lacking any sense of local knowledge. These models are also untested in those languages, so even if data are in the training set, the quality of the model may be practically unusable. 

And it’s not just Colombia. Take the current Ebola outbreak in DRC – we dug into this too, and the failure turned out to be stranger than plain absence. Congolese Swahili, the language much of the response is actually conducted in, isn’t missing from the big AI vendors’ language claims – Anthropic, OpenAI and Google all list ‘Swahili’ as supported.

But that’s the coastal variety standardised in Kenya and Tanzania, a different language under the same umbrella name, and nothing in those claims says which one you’re getting. So a responder checking whether their tools cover the language they work in gets a correct-looking answer about the wrong language. 

So we’re having to figure out: do we need to fine-tune models and collect data with teams on the ground, in this emergency context, just to improve the state of the art for those languages right now. 

So many languages, especially Indigenous languages, don’t have written text – and that’s primarily what these large language models have been trained on. English completely dominates the internet, and that’s where large language models have overwhelmingly pulled their resources from. 

You’ll have far more content in English, and for Western-dominated contexts, than you would even for Spanish or Arabic – languages you’d think of as major world languages.

And when you get down to an Indigenous language with a small population and almost no presence online, sometimes not even a written form at all, that gap goes from many times more to millions of times more, or no data at all.” 

Research theme convergence and the challenges of shadow AI

Joshua highlighted convergence in his research findings and the HLA/DFS study.

“Two of the very prominent themes that came out actually aligned very well with the research you had done at the HLA with Data Friendly Space: AI literacy in general was very low. Very few people I spoke to had any kind of formal training on how to use AI responsibly.  

Only one of the seven practitioners’ organisations had a comprehensive AI policy in place – and that organisation was shut down shortly after because of the funding cuts from US donors. It wasn’t necessarily surprising to my research, I expected that going in.

But it was quite shocking to hear directly and experience it, and a strong confirmation, through independent research, of what your team had done. 

I think it also leads into one of the findings you had with the HLA/DFS study – this idea of shadow AI use. There’s this big policy discussion happening globally, but then there’s the operational reality: these tools are helping me do my work, I don’t have enough people to help me, so I’m just going to do what works and not talk about it, because I don’t want to get in trouble.  

If we’re wanting to constrain what people can use – you can’t use this tool, you can’t use that tool – but we don’t provide an alternative that meets the current need, then we’re just driving adoption into the shadows.  
 
And once it’s in the shadows, there’s no way to check the reasoning behind any decision, because it was someone’s personal ChatGPT account all along.”

“The risk of doing nothing” 

Asked for the one message he’d want the sector to take from his research, Joshua returned to the closing line of his thesis. 

“The risks of introducing AI into humanitarian contexts are high and should not be understated, brushed over, or ignored. However, given the current state of the humanitarian sector, and the potential AI and LLMs offer to address the scale of humanitarian need,  I ultimately agree with the sentiment of two of my interviewees, who put it as a question: isn’t inaction the biggest risk? 

This is just such a big problem, we don’t really know how to address it, so let’s see what happens, and maybe someone else will solve it for us. But that’s not really a viable option right now. The risk of doing nothing is far greater than actually addressing this head-on…” 

It’s a conclusion that cuts against an assumption Joshua expected to confirm. He expected concerns about AI widening the power gap between Global North toolmakers and Global South users to dominate the conversation. Some AI experts did raise this – one named the consolidation of power in technologists’ hands as the biggest risk – but a different, more surprising thread also emerged – one that links directly to notions of shifting power and the localisation agenda: 

“A lot of the sentiment from the practitioners was that it was enabling them to do more work internally, and rely less on external organisations from the Global North to provide consulting services or advice – creating documentation internally, producing project plans and budgets and grant proposals.  

They felt like it was empowering to them, rather than putting them at a greater power disadvantage.” 

What’s next 

“Right now, as the humanitarian team lead, we’re figuring out how to use AI responsibly in emergency response, where you have very little time to produce critical information and get services to where they’re needed.  

There are things like rapid dashboard tools we trialled in the Venezuela response, and tools we’re exploring for the DRC Ebola response – transcription, translation and analysis, so that needs assessments done in audio or written form can be processed quickly enough that project plans get developed effectively and delivery of aid and services is efficient and effective. 

I’m also bringing some of my research internally into Kobo, to help lead our AI policy work, developing new tooling (LLM-driven documentation translations, support chatbots, and more) and thinking about how we bridge the same divide I saw in my research – between our technical team and our more practitioner-style team – so that each understands more of the challenges of the other. 

HLA and DFS have been pioneers in the early research of data and AI in humanitarian contexts, and we need so much more of that. There are so many gaps – in the literature, in understanding how AI is being used and how to use it responsibly. We need more collaboration between organisations willing to take on that same coordinating role.”

 

A man with short, curly brown hair and a beard stands outdoors, looking at the camera. The background is slightly blurred, showing building structures and sunlight.
Isn’t inaction the biggest risk? This is just such a big problem, we don’t really know how to address it… That’s not a viable option right now. The risk of doing nothing is far greater than actually addressing this head-on.
Joshua Beretta

Congratulations to Joshua on successful completion of his research and we wish him all the best with next steps integrating these research insights into the work of Kobo. This interview was conducted to build on and highlight the impact of the Humanitarian Leadership Academy and Data Friendly Space Humanitarian AI 2025–26 research initiative. Supporting resources – including reports, podcasts, webinars and microlearning guides – are available on the research landing page.
 

kobo.ngo

kobotoolbox.org

Access Joshua’s thesis, “Responsible Humanitarian AI: Shared Principles, Divergent Practice”, on his personal website.

You may also be interested in a webinar held by the HLA in partnership with NetHope in July 2026: Collective action on localised humanitarian AI: Unlocking solutions together, featuring Tino Kreutzer, Chief Operating Officer at Kobo, presenting his organisation’s work on a panel discussion hosted by Ka Man Parkinson.


Disclaimer 

The views and opinions expressed in this interview are those of the featured individual and do not necessarily reflect those of their affiliated organisations. This resource has been produced as a contribution to ongoing discussions on the use of AI in the humanitarian sector. Publication does not constitute endorsement of any specific technology, individual, organisation or approach. 

Newsletter sign up