News

20250325_074733

Authors’ note: This article was written by Dr. Matthew Heaton (NISD), Dr. Susanne Dreisigacker (CIMMYT), Dr. Jesus Quiroz-Chavez (JIC) and Dr. Francisco Fernando Delgadillo Galvan (CIMMYT). .  

Researchers from the UK-CGIAR Centre’s wheat improvement project are making wheat genetic databases more accessible for Artificial Intelligence (AI) to unlock a new era in crop breeding. The research builds on the growing volume of multi-layer wheat genetics information in improved and heritage wheat varieties and the untapped traits they contain which improve wheat production. Enabling the integration of these database architectures for seamless multi-modal AI-driven approaches has the potential to directly address long-standing bottlenecks in gene discovery and functional characterization. This could improve the identification of key traits and accelerate their introgression into improved wheat varieties for better human and planetary health.

The wheat genome is nearly six times larger than the human genome and at one point was thought to be impossible to sequence for research use. Today, the situation is radically different. Breakthroughs in genetic sequencing technology mean that large wheat collections are being sequenced faster than researchers can act on the information provided.

For instance, the Watkins Collection of heritage wheat plants contains data on over 800 lines, including a huge range of new unknown traits and genes not found in modern wheat varieties. The challenge now is how best to understand the treasure trove of information in these datasets and how to help wheat breeders develop better varieties.

Dr Susanne Dreisigacker of the International Maize and Wheat Improvement Center (CIMMYT) said, “wheat research is advancing at an incredible pace and AI offers further potential here.  New AI-driven automated workflows will help further streamline the translation of complex sequence information into precise, field-ready breeding decisions”

AI offers huge potential to identify key traits from wheat databases, such as the Watkins Collection, but this data is not optimised for AI analysis. Currently wheat genetic information is held in the form of short genomic sequences or assemblies which need to be rendered into consistent and comparable structures to work best for machine learning models. The project research teams from CIMMYT therefore trialled an approach which involved mining the Watkins Collection together with CIMMYT and ICARDA’s digital gene bank information for useful genetic diversity. This information will be used to pave the way for future AI-driven automated workflows.

The team used ‘k-mers’, which are short substrings of DNA sequences, to detect genetic variation across thousands of wheat accessions. Unlike other approaches to mapping genetic variation, k-mers allow direct comparisons across genetically diverse plants, without relying on a single reference genome – making them ideal for use with more distantly related wild relatives. The system is also more straightforward to implement, providing a more scalable approach with less computational runtime. Using this system also offers a more effective framework for exploiting large-scale sequence datasets in breeding.

Dr Jesus Quiroz-Chavez of JIC-CIMMYT said, “the k-merapproach is already helping us to identify valuable genome loci. Building on these tools, we’re developing a system to accelerate and scale trait discovery,  thus making the genetic data more accessible and useful for researchers and breeders.”

The k-mer approach has been used to develop a new landrace core collection –  enriched to identify potentially valuable genetics in wild relatives –  which has initially been evaluated for  resistance against fungal diseases, like wheat yellow rust (Puccinia striiformis). 

Yellow rust is a major threat to wheat harvests globally. Growing genetically resistant wheat varieties is the primary defence for farmers against the crop disease. Like new flu strains, however, the rust fungus continues to change, bringing with it the potential to overcome previously resistant varieties. Staying ahead of the rust depends upon continuing to breed resistant varieties.

Novel yellow rust resistance genes have already been identified in the Watkins Collection. Further investigation by the CIMMYT team’s discovery has confirmed the identified genes do not exist in CIMMYT’s modern wheat varieties. Thanks to this work, these new genetics can be incorporated into breeding pools for researchers across the globe to use, providing essential new resistance breeding targets to continue to protect farmers’ harvests against rust outbreaks.

Fernando Galvan from CIMMYT said, “identifying these new rust resistance genes is just one small example of the huge potential within sequenced collections. By combining the latest lab approaches with AI, we will see many more such discoveries and a much faster way to bring these to the field.”

Developing rust resistance is just the first step and this same pipeline is already being applied to other wheat breeding targets. The new AI-readiness approach developed by the project team continues to feed into global wheat research and guide how AI-driven automated breeding workflows can ultimately lead to more resilient, productive, nutritious and sustainable varieties. 

Dr Quiroz-Chavez said “genomic information is now easy and cheap to generate, including for wheat across large genebank collections. The big challenge, however, is to store, transfer and analyse this data due to computer constraints. AI tools could leverage this vast amount of data for better predictions to accelerate genetic gains and breeder applications. We now need to invest in a diversity of use cases of deployment and computer infrastructure to take full advantage of the potential of AI-enhanced crop breeding.”

Image credit: Susanne Dreisigacker.