Using machine learning to predict plants’ genetic ‘starting line’
NSF-funded project aims to determine the precise location transcription begins across plant kingdom, from maize to moss
Like a finely choreographed performance, the journey from genes to living organism is one of perfect positioning and impeccable timing. Genetic instructions provided by DNA are transcribed into messenger RNA, before being translated into the proteins necessary for cellular functions.
For decades, plant scientists haven’t had a reliable and timely way to predict where the “starting line” for this one-way process of gene expression begins.
Now, with a new $1.7 million grant from the National Science Foundation, or NSF, a team of researchers with expertise in genetics, molecular biology and artificial intelligence is looking to change that.
Comprised of Michigan State University's Erich Grotewold and Andrea Doseff, and in collaboration with Molly Megraw of Oregon State University, the scientists are aiming to build a tool they’re calling BioLearnTSS.
A machine learning framework, BioLearnTSS will help predict the precise location that genetic transcription kicks off across the plant kingdom, from maize to moss, simply by reading DNA sequences — a breakthrough that aims to reshape approaches to plant science.
“While this is fundamental research, it creates the knowledge base needed to more precisely improve crops with greater resilience, enhanced nutritional quality and increased production of health-promoting compounds,” said Doseff, a professor in the Department of Physiology as well as Pharmacology and Toxicology.
On your mark, get set – go!
BioLearnTSS gets its name from the transcription start site, or TSS. If DNA were a track, the TSS would be the starting line for a race run by molecular machines called RNA polymerases — enzymes responsible for synthesizing messenger RNA.
Just upstream of the TSS, you’ll find what’s known as the core promoter. This segment acts as a staging area of sorts for RNA polymerase and other proteins called transcription factors that ultimately help decide when, where and how much of a gene is going to be expressed.
When researchers want to map particular TSSs today, they’re looking at an intensive, species-by-species process.
"We can annotate protein-coding regions of a genome very easily, but there's no way to predict, just from sequence, where the messenger RNA actually starts," said Grotewold, an MSU Research Foundation Distinguished Professor in the Department of Biochemistry and Molecular Biology.
"That creates a major bottleneck."
To complicate matters, genes frequently use more than one TSS. Like runners jumping into a race at different points around the track, these alternative transcription start sites, or aTSSs, can lead to a wide range of outcomes.
“Transcripts produced from different TSSs could be regulated differently, and even produce functionally distinct proteins,” said Ashton Datko, a graduate student in the Grotewold Group who’s currently investigating the relationship between core promoters in maize and a gene’s TSS usage.
Genes in the machine
To accurately predict their start sites, the team will train computational models on existing TSS data from flowering plants and then test how well these models map TSSs in conifers, mosses and green algae.
That training means teaching a computer to recognize the same patterns RNA polymerase itself relies on in nature when picking where to begin DNA transcription.
“You might think of RNA polymerase as a hungry molecular machine whose job is to decide which patterns of ‘protein lollipops’ it likes along that DNA well enough to hang out and begin making transcripts there,” said Megraw, a computational biologist who specializes in machine learning at Oregon State University's departments of Botany and Plant Pathology and Computer Science.
“It’s an amazing molecular pattern recognizer, and we want to know how it does its job,” she added.
Using CRISPR gene editing in maize and the model organism Arabidopsis, the researchers will also test directly which DNA sequences control where transcription begins. Collectively, the project’s findings aim to get new, powerful tools into the hands of plant scientists worldwide.
"Understanding how transcription factors read the genome to select different start sites really hasn't been explored in almost any system," said Grotewold, who also holds an appointment in MSU’s Department of Plant Biology.
“Knowing exactly where these starting points are and how they are controlled is essential for understanding how plants grow, respond to their environment and produce traits that matter for agriculture and human health,” echoed Doseff, who also emphasized the importance of interdisciplinary teamwork when tackling sweeping scientific challenges.
The Doseff and Grotewold Groups have collaborated on numerous projects over the last two decades, and with its fusion of biochemistry, machine learning and genetics, the latest undertaking with Megraw is further proof of the edge gained when bringing together an array of expertise.
“We've published predictive modeling work in this vein of Arabidopsis, and I'm excited to take it forward in more complex genomes with Erich and Andrea, who are world-class experts in plant biochemistry and molecular biology,” Megraw said.
- Categories: