When engineering biology, there are genes, and then there are complex genes. Gene synthesis is a foundation underlying much of the experimentation within biology, but many designed DNA sequences are laden with features that make them difficult, if not impossible to reliably synthesize. Many solutions that offer complex synthesis are still shrouded in timeline and production uncertainty, making experimental planning a headache.
On the other hand, advances in AI and genetic engineering are driving a fast-growing need for the reliable synthesis of complex genes. Protein engineers using AI often want to synthesize their models’ exact output without the need for further optimization, ensuring minimal extraneous factors in their design-build-test-learn cycle. Similarly, wresting control over gene expression—as would be needed for recombinant proteins or genetic circuits—is likely to require promoters, enhancers, and other complex regulatory sequences that can’t be optimized without damaging their function. When a project needs sequences that can’t be optimized or changed, the need for complex sequences leads to uncertainty and increased risk.
To help you mitigate these challenges, Twist Bioscience has brought its expertise in oligonucleotide synthesis and automation to bear on the problem; developing the means to faithfully produce 99.5% of researcher-requested genes within 15 business days of ordering, all while maintaining full adherence to biosecurity screening and regulatory guidelines. But don’t just take our word for it. Below, we delve into the meaning of complex genes, using fulfilled customer orders as real-world examples.
A complex gene is any DNA sequence whose collective features make it difficult to produce synthetically. Synthesis companies will have different complexity scoring and acceptance criteria, therefore the definition of a complex gene will vary by company. If a sequence is rejected by synthesis providers, it likely has to do with the following features:
Collectively, these challenges mean that synthesis companies must either reject complex sequences (those with high GC content or repetitive elements), or else they must develop bespoke solutions to both synthesize and quaility control (QC) them. Twist set out to solve this problem. Leveraging ultra-long, silicon-based oligo synthesis, automated enzymatic assembly, and long-read sequencing, Twist has developed the ability to reliably synthesize complex genes. Not only does this new approach greatly expand the breadth of sequences that can be accepted for synthesis, but it enables their production within 15 business days**.
Already, this method has proven capable of manufacturing genes that would likely be rejected from most other platforms.
🤔 Will your sequence be accepted?
Want to know if your sequence will be rejected by Twist’s synthesis platform? Click here to use our quick-check manufacturability analyzer, which provides a rapid assessment of your sequence’s complexity, an estimated cost, as well as a predicted turnaround time for its synthesis*.
High GC content can have significant influence over the complexity of the gene synthesis process. Extremely high GC content, as well as rapid swings from high to low, or vice versa, can stall polymerases and lead to secondary structure formation (specifically G-quadruplexes). Such barriers are among the most challenging for gene synthesis platforms.
If you’re studying gene regulation, for example, it is difficult to avoid GC-rich sequences. Promoters and other upstream regulatory sequences are often enriched for GC content, likely to enable epigenetic control over gene expression1,2. Whether it's to study their function or wield their influence, you need the ability to synthesize these regulatory elements without worrying about their GC content.
Twist’s Complex Gene Synthesis platform enables access to these critical genomic features, as demonstrated in Figure 1. Therein, two genes are presented: one (A) with an extremely high average GC content of 77%, and another (B) with extreme localized GC content followed by a rapid swing from high (96%) to low (15%) in 50bp windows.
Twist’s platform is able to reliably generate sequences with an average GC content between 25% and 75%, as well as localized GC content between 10% and 90% (within a 50 bp window). Notably, the genes presented in Figure 1 fall outside of these limitations, but are presented to highlight the capabilities of the platform beyond its conservative boundaries.
Figure 1: High GC Content Genes. Both graphs depict the GC percentage in a sliding 25bp window across each genes' length. The gene depicted in (A) demonstrates a staggeringly high average GC content of 77%, with a peak at 94%. The gene depicted in (B) illustrates extreme localized GC content, with a steep peak and decline in GC content ranging from 96% to 20% seen around the 1250bp mark. Both genes were successfully synthesized by Twist's complex genes platform in less than 15 business days**.
The human genome is awash with repetitive sequences, comprising roughly 50% of the genomic landscape3. Whether it’s long direct repeats, inverted repeats, short tandem repeats, or mononucleotide repeats, these elements are believed to play important roles in regulating gene expression and genome stability4,5.
Despite their importance, synthesizing genes that contain repetitive elements can be a significant challenge, in part because they’re difficult to clone, amplify and sequence. With techniques like solid-phase phosphoramidite chemistry, long and repetitive sequences can be faithfully synthesized and assembled. But these nascent assemblies must eventually be cloned to ensure sufficient material is available for the desired application.
Cloning a repetitive gene sequence is often difficult because polymerases are prone to slippage. Imperfect rejoining results in the potential growth or deletion of the repetitive sequence. Such changes are non-trivial and have the potential to significantly alter the sequence’s function4,6.
Therefore repetitive elements are a common reason for complex genes being rejected by synthesis providers. As the following examples highlight, Twist’s approach to complex gene synthesis is designed to reduce rejections, in part by overcoming the barrier of repetition.
Eukaryotic genomes are peppered with short stretches of mononucleotide repeats (also known as homopolymers). Homopolymers are suspected to have varied roles in genomic function, whether it's altering the structure of the DNA helix or providing a landing spot for effector proteins7. With roughly 1.43 million homopolymers in the human genome, these short satellite repetitive sequences have to be accounted for in genomic studies8.
Yet, polymerase slippage and sequencing challenges make homopolymers particularly challenging to synthesize. For this reason, Twist’s platform was designed to make homopolymers more accessible, enabling the synthesis of mononucleotide sequences up to 30 base pairs in length. Figure 2 shows a gene roughly 700bp in length that was recently synthesized by Twist’s platform. This sequence would be classified as highly complex owing to a high average GC content (~67%) and a poly-A stretch that spans 30 base pairs.
Figure 2: Homopolymer Synthesis. (A) Graph depicting the GC percentage in a sliding 25bp window across the gene’s length. This gene has several complexity features, including an average GC content of 67%, inverted repeats and short tandem repeats. However, most apparent is a poly-A homopolymer stretch at the gene terminus.
Short tandem repeats are common features, both in natural genomes and in synthetic biology. One increasingly popular example is the glycine-serine linker, a repetitive peptide bridge that helps to connect two disparate protein domains 9, as in the joining of heavy and light chain domains in single-chain variable fragments (scFv) 10.
Such a linker is critical for the formation of fusion proteins, but it is also a challenge to synthesize. Not only can polymerases slip from the repetitive sequence, but the rapid succession of guanine nucleotides can lead to the formation of G-quadruplexes. These secondary structures can disrupt amplification and are prone to cleavage, both of which may lead to failed gene synthesis. Accordingly, the challenge of tandem repeat synthesis is one of the primary reasons that fusion-protein gene sequences are rejected.
Twist’s complex gene synthesis platform enables the formation of short tandem repeats consisting of 3-9 base pairs repeated up to 100 total base pairs. This means Twist’s platform can faithfully synthesize glycine-serine peptide linkers. Figure 3 shows data from a fusion-protein gene synthesized on Twist’s complex genes platform. A linker region in the middle of the gene consists of a 15 base pair unit that repeats 13 times, creating a sequence that is very prone to G-quadruplex formation. Nonetheless, Twist was able to reliably produce this gene in under 15 business days**.
Twist’s ability to synthesize such a sequence means that fusion-protein genes with linker-heavy designs can be reliably produced on a predictable timeline.
Figure 3: Short Tandem Repeat Synthesis. (A) Graph depicting the GC percentage in a sliding 25bp window across the gene’s length. Evident in the pattern is a region in the middle of the gene, where a 13x tandem repeat made up of 15bp spans 195 base pairs. While this gene exceeds the conservative boundaries of Twist’s platform, it was nonetheless faithfully synthesized, demonstrating the far ranging capabilities of complex gene synthesis.
Larger repeats come with unique challenges. Long repeats, for example, are prone to recombination-driven errors that can result in deletions or unintended fusions during cloning. Inverted repeats, on the other hand, are highly prone to forming obstructive secondary structures, particularly hairpins. For these reasons, both long and inverted repeats are difficult to reliably produce without bespoke solutions.
As with homopolymers and short tandem repeats, Twist’s platform is capable of consistently generating sequences with direct long repeats (wherein a sequence up to 200 base pairs is repeated) as well as inverted repeats (up to 100 base pairs repeated). Long repeats can be seen in Figures 3 and 4, representing two different genes successfully produced on Twist’s platform. Figures 3 and 5 similarly show sequences for synthesized genes containing inverted repeats.
Figure 4: Synthesizing Long Repeats. Graph depicting the GC percentage in a sliding 25bp window across gene’s length. In the graph, the gene’s highly repetitive nature can be readily seen. Despite the complexity of this repetition, synthesis start to packaging finished in just 13 calendar days**.
Figure 5: Synthesis of Inverted Repeats. Graph depicting the GC percentage in a sliding 25bp window across gene’s length. In addition to extreme GCs content swings (from 20% - 90%) this gene also contains two inverted repeats at the gene’s termini, representing common AAV vector elements. Here too, synthesis start to packaging finished in just 13 calendar days**.
With Twist, you are no longer left in the dark about the process and progress of complex gene synthesis. When building the product, many researchers told us that it's common to place an order with a provider that claims to accept everything, only to receive little to no communication about when their order will be manufactured and delivered. Sometimes it could take a few weeks for synthesis providers to fulfill the order. Other times, it could take months. We set out to make it significantly easier to plan experiments around complex gene sequences, especially when on a tight timeline and budget.
Twist’s Complex Gene Synthesis platform provides a transparent, easy ordering process. Using your Twist account, you can receive an immediate price quote and estimated turnaround time before your order is placed. Once it is being made, you can live track the manufacturing process. When rare delays are encountered, Twist’s DNA experts immediately get in touch.
In other words, the genes may be complex, but everything else is made simple. This translates to a more reliable and transparent synthesis process for the hardest to make genes. Repetitive sequences, high GC content, and vulnerability to secondary structures used to delay or block research. But with Twist, these complex genes can become simple and routine.
*Pricing and turnaround estimates subject to final sequence validation
**Turnaround time may vary based on sequence parameters and biosecurity screening.