How Can We Design Better AI For Agriculture Using Better Data?
By Jawoo Koo, Senior Research Fellow, IFPRI for CGIAR Digital Transformation Accelerator
31 July 2026, US: From predicting crop performance and identifying pests to supporting breeding decisions and helping farmers adapt to climate change, AI is opening up new possibility for agriculture.
But there is one obstacle that no algorithm can overcome on its own: data.
AI models continue to grow in size and sophistication while many of the world’s agricultural data gaps remain unresolved, particularly in the Global South, where better data could have the greatest impact.
Bringing agricultural knowledge together
The good news is that agriculture already has a wealth of publicly available information. Through years of collaboration with partners, CGIAR has developed GARDIAN, a platform that brings together agricultural publications, databases and datasets in one place. Today, GARDIAN indexes more than 470,000 publications and digital assets, 26,000 datasets and over half a million standardized agronomic data points, making it easier for researchers and AI developers to discover and use agricultural knowledge.
But making data available and standardized is only part of the challenge.
The danger of looking where the light is
A few years ago, we worked on a project to understand how market access influenced farmers’ adoption of agricultural technologies. We used a mobile phone survey to collect data from a country in East Africa.
Initially, the results made sense: the further farmers were from markets, the lower the adoption of new technologies. But then we noticed something unexpected. In the most remote areas, technology adoption appeared to increase sharply again.
At first, we wondered whether these farmers were simply more committed to agriculture and therefore more willing to adopt new technologies.
The reality was much simpler: selection bias.
The only people who could afford mobile phones in those remote areas were generally wealthier farmers. They were also the ones who could afford improved technologies. Our survey had unintentionally excluded poorer households.
This selection bias was caused by the streetlight effect: looking for your lost keys under the streetlight because that’s where the light is, rather than where you actually dropped them. By only surveying farmers who were better off (under the streetlight), we ended up excluding more disadvantageous farmers (leaving them in the dark).

Agricultural data faces a similar challenge. There are still many “dark corners” across the agricultural landscape, particularly in the Global South, where collecting quality, representative data remains difficult.
AI can only learn from the data we give it, so if some communities or farming systems are missing, those blind spots become part of the model and result in bias.
Making agricultural data visible to AI
Even when valuable agricultural data exists, AI models may never use it.
In 2023, The Washington Post examined the sources used to train leading AI models. CGIAR’s research outputs were barely represented.

It wasn’t fully clear why, but we implemented multiple strategies to address this.
First, we reprocessed CGIAR’s knowledge in AI-friendly formats and published through Hugging Face, one of the world’s leading platforms for AI developers and researchers. This helps make agricultural knowledge — and, in particular research from the Global South — more visible to the AI community.
Licensing is another challenge. Most CGIAR content is published under the Creative Commons Attribution license (CC BY), which requires attribution when content is reused. While this works well for people in the agricultural research community; it is much harder for machines and AI models to provide attribution at scale.
To bridge that gap, CGIAR is partnering with CABI to develop a Model Content License that allows research outputs to be used for AI training while preventing models from reproducing the original content verbatim. The approach is currently being tested with publishers that have proprietary content.
Global biases remain deeply embedded
Improving accessibility does not automatically eliminate bias.
A 2023 study analyzing more than six million scientific publications highlighted persistent biases inequalities in research production, with much of it originating from just 10 high-income countries.

The same patterns appear in geospatial datasets.
Take OlmoEarth, an AI foundation model built from hundreds of thousands of Earth observation data points, as an example. When you map the available data used to develop the model, Africa is clearly underrepresented. This is not because OlmoEarth ignores African data: it simply reflects what has been published and made available in the geospatial research community over the years.

Farming systems in the Global South are fundamentally different from those in high-income countries. Farm sizes, cropping systems and management practices vary enormously. AI models trained primarily on data from Europe or North America risk producing biased predictions and recommendations when applied elsewhere.

Rather than assuming that larger models are always better, CGIAR has found that locally trained, context-specific models can often deliver stronger results. Collaborating on a crop-mapping initiative in Kenya with Olmo Earth, for example, demonstrated how context-specific data can improve AI performance.
Speaking the same language
Interoperability is another challenge.
Different systems may describe the same information in different ways.
Within CGIAR, the Enterprise Breeding System (EBS) supports harmonized data collection on research stations for breeding programs, while ClimMob captures the performance of improved crop varieties in farmers’ fields.
Although the two systems contain complementary information, they have historically described crop varieties differently. Some inconsistencies were easy to fix, such as typos or punctuation. Others reflected years of different practices across breeding teams.
For soybean breeding alone, we found that around half of the records were initially not interoperable because different teams had recorded the same varieties using different naming conventions.
These challenges are not unique to CGIAR; they are often organizational and cultural. Without interoperability, even the most advanced AI systems struggle to connect the dots.
Where do we go from here?
Building better AI for agriculture starts with building better data.
That means continuing to fill known data gaps, increasing the visibility of research from the Global South, designing interoperable systems from the outset, and making agricultural knowledge more accessible for responsible AI development.
Most importantly, it means keeping people in the loop. AI can help us discover patterns and generate insights, but researchers, breeders and farmers remain essential for interpreting results, identifying bias and ensuring that AI reflects the realities of agriculture.
Because when we fix the data, we improve everything that comes after it.
Also Read: Bayer Launches New Insecticide Trance for Cotton in India
Global Agriculture is an independent international media platform covering agri-business, policy, technology, and sustainability. For editorial collaborations, thought leadership, and strategic communications, write to pr@global-agriculture.com






