
OpenCRISPR-1: Generative AI Meets CRISPR
By: Christos Evangelou
Researchers used large language models (LLMs) to expand the sequence diversity of CRISPR-Cas proteins and generate novel, functional genome editors with improved properties compared to natural systems. They publicly released OpenCRISPR-1, the first AI-designed editor for precise genome editing. This novel AI-based approach for designing CRISPR gene editors could potentially expand the capabilities and applications of genome-editing technologies.
In a recent study, researchers at Profluent Bio (Berkeley, CA, USA) used artificial intelligence (AI) to design novel CRISPR gene editors that are functional and that show comparable or improved genome-editing activity and specificity relative to naturally occurring gene editors, despite being hundreds of mutations away from any known natural protein. This innovative approach has yielded OpenCRISPR-1, the first AI-generated gene editor.
In a proof-of-concept study, OpenCRISPR-1 demonstrated comparable efficiency to the widely used Streptococcus pyogenes Cas9 (SpCas9) while offering improved specificity. This development not only expands the CRISPR toolbox but also paves the way for creating gene editors tailored to specific applications, which could range from agriculture to medicine.
»Our LLMs, trained on billions of proteins in nature, are able to learn the sequence-to-function mapping of natural proteins and can be utilised to build functional proteins from scratch, such as OpenCRISPR-1,« said Ali Madani, PhD, founder and CEO of Profluent Bio and senior author of the study.
»We can now use the model as a guide to install additional desired features, for example, a set of mutations that alter PAM selectivity, a combination of deletions to reduce size, or a set of mutations to alter thermostability while simultaneously maintaining other key properties, such as processivity and stability,« he added. »This can be very difficult to achieve with directed evolution because the addition of one property may impair another. AI can circumvent this.«
The study has not been peer-reviewed and is available as a preprint on bioRxiv.
Rationale: Addressing limitations of natural CRISPR-Cas systems
CRISPR-Cas systems, originally evolved as bacterial defence mechanisms against viruses, have been repurposed as powerful gene-editing tools. However, these natural systems often face limitations when applied in non-native environments, such as human cells.
Traditional approaches to optimising these tools for use in non-native environments include directed evolution and structure-guided mutagenesis. Although these approaches have yielded improvements, they are limited by the need for explicit structural hypotheses or complex screening processes.
Directed evolution, for instance, can be limited by the complex nature of ‘fitness landscape’, a way to visualise how different genetic variations perform. On the other hand, structure-guided approaches depend on solved structures representing key functional states that are often difficult to obtain for complex functions beyond simple binding interactions.
The objective of this study was to edit the human genome for the first time using a gene editor that was designed entirely using AI. »This was a scientific moon-shot, as CRISPR-Cas proteins are incredibly complex molecular machines with precise functions that require an understanding of protein-protein as well as protein-nucleic acid interactions,« said Dr. Madani.
Approach: Using AI to design novel, optimised editors
To overcome the limitations of traditional approaches for optimising gene editors, the research team leveraged the power of large language models (LLMs) trained on vast amounts of biological data to generate novel CRISPR-Cas proteins that could function as efficient gene editors in human cells.
To this end, the team compiled the CRISPR-Cas Atlas, an extensive dataset of over one million CRISPR operons from diverse microbial genomes, by mining 26 terabases of assembled genomes and metagenomes.
»To our knowledge, CRISPR-Cas Atlas is the most extensive dataset of CRISPR systems curated to date,« emphasised Dr. Madani.
Using this dataset, the researchers fine-tuned ProGen2, a protein language model previously developed by Profluent, to specialise in generating CRISPR-Cas proteins. The team balanced the training data for protein family representation and sequence cluster size to ensure broad coverage. The fine-tuned models were used to generate four million novel CRISPR-Cas protein sequences. Half were generated directly, while the other half were prompted with short segments from natural proteins to guide generation towards specific families.





