A powerful new foundation for custom queries—built on Lucene and designed for R&D precision.
Over the past few years, Cypris has helped innovation teams make faster, more informed decisions by centralizing critical insights across datasets like patents, academic papers, and company activity. But until now, our search experience relied on a legacy query system with limited capabilities, offering little support for advanced search features or dataset-level customization.
Today, we’re excited to introduce an upgraded Advanced Search on Cypris, a complete overhaul of our query engine and search experience, powered by the open-standard Lucene query syntax. This update introduces a more robust and flexible search foundation, unlocking new ways to query data, build complex filters, and extract precisely what you need across patents, research, and more.
Why we rebuilt our search system from the ground up
Cypris’ original query syntax, a proprietary format used internally for years, limited users’ ability to craft advanced queries or tailor searches to specific datasets. It lacked modern capabilities like proximity searches, field-level customization, or true Boolean logic. This made it difficult to build a reliable and intuitive experience for both casual users and advanced researchers.
By moving to Lucene, we’re adopting a powerful, industry-standard query language that makes it easier for developers to build advanced features—and gives users access to a far more capable and flexible search toolset.
What’s new in Advanced Search
1. Custom Queries by Dataset
You can now layer queries to search across datasets or tailor filters to each one. For example, you can run a broad query on drone delivery, and then add separate layers to focus on patents by a specific assignee and papers from a specific country or funding agency.
Navigating the All Datasets tab introduces a new level of complexity—and power—by allowing users to apply dataset-specific logic within a single, unified query workflow. While querying multiple datasets simultaneously might seem straightforward, the underlying differences in schema, metadata, and available fields between our proprietary datasets make this a deeply technical challenge. Patents, for example, include claims, application numbers, and multiple date fields (filed, granted, updated), while academic papers use DOIs, have different structural conventions, and emphasize different metadata. In the past, we sidestepped this complexity by translating general queries like ((drone_allText)) into dataset-specific logic under the hood. Now, instead of obscuring that logic, we allow users to opt in to it. The builder provides progressive layers of customization: start with intuitive keyword searches across all fields, then move into the advanced builder for field-specific targeting, fuzzy logic, and term boosting, and finally, tailor query logic by dataset—such as specifying different countries of interest for papers vs. patents. This approach preserves flexibility while giving users full control, and with tools like our real-time Live Analysis and “Your Query” panel, we make it easy to understand how every decision affects the results.
2. More Fields to Query
We’re exposing deeper fields across datasets—giving you explicit control over the dimensions of your search. For the first time, users can now search academic papers by DOI, a critical identifier previously unsupported on the platform. You can also query by:
- Author or inventor names
- Organizations or assignees
- Countries, journals, funding agencies, and more
3. Full Boolean Support
Advanced Search now leverages powerful Boolean logic—AND, OR, NOT, and grouping—enabling more precise control over search logic and improving performance and accuracy.
4. Lucene Syntax Features
Use built-in Lucene features to create expressive, complex searches:
- Proximity searches to find terms near each other
- Fuzzy searches for flexible matching
- Exact phrase matching
- Boosting to prioritize results (e.g., prioritize results mentioning AI 3x more than others)
- Prefix/Postfix queries to match phrases that start or end a certain way
- Range queries for fields like date, funding amounts, or numerical values
A more powerful user experience
Our new search interface is built to help you tap into these capabilities without needing to know the syntax from the start. You’ll find:
- A Query Builder to guide you through complex searches
- A Help Video to onboard users to Lucene-style searches
- Inline examples and tips for writing queries using grouping, boosting, and more
Built for precision, speed, and customization
With Lucene as our foundation, search results are now not only more flexible but also faster and more accurate. Semantic search continues to offer natural-language ease of use, while Boolean search gives power users the performance and structure they need to uncover insights with greater specificity.
Whether you’re an innovation analyst drilling into AI patents or a business development lead scanning academic papers from Chilean researchers—Advanced Search is built to help you get to the signal, faster.
Available now to all users
Advanced Search is live and available across the Cypris platform today. If you’re already using Cypris, you’ll find the new search interface in your dashboard, complete with updated syntax documentation and walkthroughs.
We’re excited to see what you’ll build, discover, and analyze with this new capability. This is just the beginning—we’ll continue expanding the fields, syntax features, and customization options as we push the boundaries of what intelligent search can do for R&D.

Introducing Advanced Search on Cypris

A powerful new foundation for custom queries—built on Lucene and designed for R&D precision.
Over the past few years, Cypris has helped innovation teams make faster, more informed decisions by centralizing critical insights across datasets like patents, academic papers, and company activity. But until now, our search experience relied on a legacy query system with limited capabilities, offering little support for advanced search features or dataset-level customization.
Today, we’re excited to introduce an upgraded Advanced Search on Cypris, a complete overhaul of our query engine and search experience, powered by the open-standard Lucene query syntax. This update introduces a more robust and flexible search foundation, unlocking new ways to query data, build complex filters, and extract precisely what you need across patents, research, and more.
Why we rebuilt our search system from the ground up
Cypris’ original query syntax, a proprietary format used internally for years, limited users’ ability to craft advanced queries or tailor searches to specific datasets. It lacked modern capabilities like proximity searches, field-level customization, or true Boolean logic. This made it difficult to build a reliable and intuitive experience for both casual users and advanced researchers.
By moving to Lucene, we’re adopting a powerful, industry-standard query language that makes it easier for developers to build advanced features—and gives users access to a far more capable and flexible search toolset.
What’s new in Advanced Search
1. Custom Queries by Dataset
You can now layer queries to search across datasets or tailor filters to each one. For example, you can run a broad query on drone delivery, and then add separate layers to focus on patents by a specific assignee and papers from a specific country or funding agency.
Navigating the All Datasets tab introduces a new level of complexity—and power—by allowing users to apply dataset-specific logic within a single, unified query workflow. While querying multiple datasets simultaneously might seem straightforward, the underlying differences in schema, metadata, and available fields between our proprietary datasets make this a deeply technical challenge. Patents, for example, include claims, application numbers, and multiple date fields (filed, granted, updated), while academic papers use DOIs, have different structural conventions, and emphasize different metadata. In the past, we sidestepped this complexity by translating general queries like ((drone_allText)) into dataset-specific logic under the hood. Now, instead of obscuring that logic, we allow users to opt in to it. The builder provides progressive layers of customization: start with intuitive keyword searches across all fields, then move into the advanced builder for field-specific targeting, fuzzy logic, and term boosting, and finally, tailor query logic by dataset—such as specifying different countries of interest for papers vs. patents. This approach preserves flexibility while giving users full control, and with tools like our real-time Live Analysis and “Your Query” panel, we make it easy to understand how every decision affects the results.
2. More Fields to Query
We’re exposing deeper fields across datasets—giving you explicit control over the dimensions of your search. For the first time, users can now search academic papers by DOI, a critical identifier previously unsupported on the platform. You can also query by:
- Author or inventor names
- Organizations or assignees
- Countries, journals, funding agencies, and more
3. Full Boolean Support
Advanced Search now leverages powerful Boolean logic—AND, OR, NOT, and grouping—enabling more precise control over search logic and improving performance and accuracy.
4. Lucene Syntax Features
Use built-in Lucene features to create expressive, complex searches:
- Proximity searches to find terms near each other
- Fuzzy searches for flexible matching
- Exact phrase matching
- Boosting to prioritize results (e.g., prioritize results mentioning AI 3x more than others)
- Prefix/Postfix queries to match phrases that start or end a certain way
- Range queries for fields like date, funding amounts, or numerical values
A more powerful user experience
Our new search interface is built to help you tap into these capabilities without needing to know the syntax from the start. You’ll find:
- A Query Builder to guide you through complex searches
- A Help Video to onboard users to Lucene-style searches
- Inline examples and tips for writing queries using grouping, boosting, and more
Built for precision, speed, and customization
With Lucene as our foundation, search results are now not only more flexible but also faster and more accurate. Semantic search continues to offer natural-language ease of use, while Boolean search gives power users the performance and structure they need to uncover insights with greater specificity.
Whether you’re an innovation analyst drilling into AI patents or a business development lead scanning academic papers from Chilean researchers—Advanced Search is built to help you get to the signal, faster.
Available now to all users
Advanced Search is live and available across the Cypris platform today. If you’re already using Cypris, you’ll find the new search interface in your dashboard, complete with updated syntax documentation and walkthroughs.
We’re excited to see what you’ll build, discover, and analyze with this new capability. This is just the beginning—we’ll continue expanding the fields, syntax features, and customization options as we push the boundaries of what intelligent search can do for R&D.

Keep Reading

Sustainable aviation fuel has moved from pilot projects to a mandated market, and its patent landscape is distinctive because SAF is not a single technology but a set of competing production routes, each with its own feedstocks, catalysts, and process chemistry. Peer-reviewed technical reviews lay out the route taxonomy: the hydroprocessed-ester-and-fatty-acid route converts waste oils and fats into jet fuel and is currently the most mature; the Fischer-Tropsch route gasifies biomass or waste into synthesis gas and rebuilds it into hydrocarbons; the alcohol-to-jet route converts ethanol or other alcohols into jet-range molecules; and the synthetic power-to-liquid route, including methanol-mediated pathways, combines captured carbon dioxide with green hydrogen to make e-fuels with no biological feedstock at all.¹,²,³,⁴,⁷ Because each route is a distinct region of patenting, freedom-to-operate and white space analysis must treat SAF as several landscapes at once, spanning feedstock pretreatment, catalysts, conversion processes, and upgrading.
The landscape is being pulled forward by regulation more directly than most. Under the European Union's ReFuelEU Aviation regulation, the sustainable share of aviation fuel supplied at EU airports rises stepwise to 70 percent by 2050, with a dedicated sub-obligation for synthetic e-fuels and an anti-tankering rule requiring airlines to uplift most of their fuel where they operate; Switzerland adopted the ReFuelEU framework from January 1, 2026.⁹ This creates both a deadline and a guaranteed market against a very large baseline, since global commercial jet-fuel demand is on the order of 100 billion gallons a year and is projected to rise substantially by 2050.¹ The near-term response has concentrated in the waste-oil route because it is the most mature,⁴,⁵ but the mandates specifically favor synthetic e-fuels in the longer term, which is steering research and filings toward the power-to-liquid route and its underlying carbon-conversion and catalysis challenges. The patent record shows this tension clearly: across the Cypris corpus of more than 500 million patents and scientific papers, the SAF space holds roughly 6,000 de-duplicated families and grew about 3.6 times between 2022 and 2024, and on an indicative basis the Fischer-Tropsch and e-fuel routes lead patent activity, ahead of hydroprocessed waste oils, with alcohol-to-jet the smallest slice, even though the waste-oil route currently leads in deployed production capacity, a divergence between where filing and where building are concentrated. The most active assignees span engine makers, refining-and-catalysis licensors, and route pure-plays, and the United States leads on geography, followed by the United Kingdom, China, France, and the Nordic producers. Because applications publish about eighteen months after filing, the most recent catalyst and e-fuel filings are under-represented (2025 and 2026 counts are partial), so the current frontier is more active than granted-patent counts suggest.
The strategic question is which route and layer to back, and the white space sits where cost and feedstock constraints are hardest. The waste-oil route is limited by feedstock availability, so its white space is narrower; the Fischer-Tropsch and alcohol-to-jet routes turn on catalyst performance and process integration;²,³,⁸ and the synthetic e-fuel route, though earliest and most expensive, is the one the mandates most favor and the one with the most open, high-value IP, particularly in the catalysts and process designs that lower the cost of converting carbon dioxide and hydrogen into jet fuel.⁶,⁷ Reading the landscape by route, feedstock, catalyst, and process, and tracking both the patents and the underlying chemistry research, is what separates a crowded region from an open one.
Where the SAF white space is
Synthetic e-fuel catalysis. Catalysts and process designs that lower the cost of converting captured carbon dioxide and green hydrogen into jet-range hydrocarbons are the most favored by mandate and among the most open, high-value targets.⁶,⁷
Alcohol-to-jet conversion. Improved catalysts and process integration for converting alcohols to jet-range molecules are an active, still-developing route.⁸
Fischer-Tropsch from waste and biomass. Gasification, syngas conditioning, and Fischer-Tropsch catalysis for waste and biomass feedstocks are a distinct, contested layer.²,³
Feedstock flexibility and pretreatment. Technologies that broaden or pretreat feedstocks, easing the supply constraint on mature routes, are a differentiated area.⁴
Process intensification and integration. Designs that integrate steps, cut energy use, and lower capital cost are where scale-up economics are decided.⁵
How AI-powered landscape and white space analysis helps
Resolving a landscape that spans several production routes, each with its own feedstocks, catalysts, and processes, under a moving regulatory timeline, requires more than keyword search. AI-powered analysis addresses this with semantic search that clusters activity by route, feedstock, catalyst, and process across varied terminology, attribution that normalizes filers to canonical entities and tracks new entrants, and continuous monitoring that keeps pace with a mandate-driven surge. Because SAF advances appear in scientific and catalysis literature before they are patented, reading both patents and literature gives the earliest signal of where scalable routes are emerging.
Where Cypris fits
Cypris runs patent landscape and white space analysis for multi-route fields such as sustainable aviation fuel across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by production route, waste-oil, Fischer-Tropsch, alcohol-to-jet, and synthetic e-fuel, and by layer, feedstock, catalyst, conversion, and upgrading, and normalizes filers to canonical entities, so a team can resolve which routes and layers are crowded and which remain open as white space, and can track new entrants as the field scales. Semantic search across patents and scientific literature connects filings to the underlying catalysis and process research, which is where SAF advances appear first. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined route over time and flags new patents and papers as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is the sustainable aviation fuel patent landscape? The sustainable aviation fuel patent landscape is the set of patents covering the several routes used to make jet fuel with lower lifecycle emissions, including hydroprocessed waste oils, Fischer-Tropsch fuels, alcohol-to-jet, and synthetic power-to-liquid e-fuels. Each route has distinct feedstocks, catalysts, and processes. It is best understood as several landscapes rather than one.
Why is regulation shaping SAF patenting? Regulation shapes SAF patenting because binding blending mandates require a rising share of sustainable aviation fuel over the coming decades, reaching 70 percent by 2050 under the EU ReFuelEU Aviation regulation, with a dedicated sub-mandate for synthetic e-fuels. Near-term activity has concentrated in the mature waste-oil route, while the mandates steer longer-term research toward e-fuels. The patent record tracks this policy pull closely.
What production routes does the SAF landscape cover? The SAF landscape covers hydroprocessed waste oils and fats, Fischer-Tropsch fuels from gasified biomass or waste, alcohol-to-jet conversion, and synthetic power-to-liquid e-fuels made from captured carbon dioxide and green hydrogen. Each is a distinct region of patenting. Freedom-to-operate and white space analysis must treat them separately.
Where is the white space in SAF? The white space sits in synthetic e-fuel catalysis, alcohol-to-jet conversion, Fischer-Tropsch from waste and biomass, feedstock flexibility and pretreatment, and process intensification. The mature waste-oil route is comparatively crowded and feedstock-limited. The most open, high-value opportunities are in the e-fuel catalysts and processes the mandates most favor.
Why is the synthetic e-fuel route strategically important? The synthetic e-fuel route is strategically important because the mandates specifically favor it in the longer term, it has no biological feedstock limit, and it is the least mature and most expensive route, which leaves the most open, high-value IP. The central challenge is lowering the cost of converting carbon dioxide and hydrogen into jet fuel. That is where much of the defensible catalysis and process IP is concentrating.
Why does the patent record differ from deployed capacity in SAF? The patent record differs from deployed capacity because filing tends to run ahead of building. In the Cypris corpus the Fischer-Tropsch and e-fuel routes lead in patent activity, even though the hydroprocessed waste-oil route currently leads in installed production capacity. That divergence signals where developers expect the next phase of growth.
What software helps analyze the sustainable aviation fuel patent landscape? Software for the SAF landscape should cluster activity by production route and process layer, resolve filers to canonical owners, search patents and scientific literature semantically, and monitor a mandate-driven field continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams use SAF patent landscape analysis? SAF patent landscape analysis is used by R&D, innovation, IP, and strategy teams at fuel producers, chemicals and catalysis companies, airlines and energy majors, and their partners, as well as investors and policymakers. It informs which route to back, where to file, and where competitors are concentrated. Cypris serves hundreds of enterprise customers across chemicals, energy, advanced materials, and other regulated industries.
Endnotes
- Heyne, J., Holladay, J., & Abdullah, Z. (2020). Sustainable aviation fuel: review of technical pathways. Pacific Northwest National Laboratory / U.S. Department of Energy, Bioenergy Technologies Office. https://doi.org/10.2172/1660415
- Zhang, X., Zheng, Y., Li, J., & Wang, X. (2025). Research advances and future perspectives in Fischer-Tropsch synthesis for sustainable aviation fuel. Sustainable Energy & Fuels. https://doi.org/10.1039/d5se01412c
- Vreugdenhil, B., Boymans, E., Viar, H., et al. (2025). Syngas to sustainable aviation fuel: emerging catalysts and routes. Applied Catalysis A: General. https://doi.org/10.1016/j.apcata.2025.120554
- Chang, K., Ng, J., Japar, W. M. A. W., et al. (2026). Lipid feedstocks for sustainable aviation fuel via HEFA: status and challenges. Renewable and Sustainable Energy Reviews. https://doi.org/10.1016/j.rser.2026.117006
- Gómez, J., & Gyandoh, D. (2025). Techno-economic analysis of HEFA and lignocellulosic biomass conversion for sustainable aviation fuel. Applied Energy. https://doi.org/10.1016/j.apenergy.2025.126421
- Riaz, A., Qyyum, M. A., Al-Muhtaseb, A. H., Al-Jahwari, F., & Saeed, A. (2026). Carbon-derived and biomass-based sustainable aviation fuel pathways: a comparative techno-economic and life-cycle review for aviation decarbonization. Carbon Capture Science & Technology. https://doi.org/10.1016/j.ccst.2026.100641
- Karlsruhe Institute of Technology (2025). Sustainable aviation fuel production via the methanol pathway: a technical review. Sustainable Energy & Fuels. https://doi.org/10.5445/ir/1000187428
- Probabilistic technoeconomic analysis of alcohol-to-jet sustainable aviation fuel: implications for design and decision making (2026). https://doi.org/10.1088/2977-3504/ae7801/v2/review1
- European Commission, Directorate-General for Mobility and Transport. ReFuelEU Aviation. https://transport.ec.europa.eu/transport-modes/air/environment/refueleu-aviation_en

Cellular reprogramming has become one of the most closely watched areas in longevity biotechnology, and its patent landscape is distinctive because the leading approach builds directly on an already foundational technology. Full reprogramming, using the four Yamanaka factors, resets an adult cell all the way to a pluripotent, embryonic-like state; partial or transient reprogramming instead applies a subset of those factors briefly, aiming to roll back the epigenetic state of an aged cell toward a younger profile while preserving its identity and function. In animal models, partial reprogramming has ameliorated age-associated hallmarks and, in one landmark study, restored youthful epigenetic patterns and recovered vision after optic-nerve injury, evidence that framed aging partly as a loss of epigenetic information that reprogramming can help reverse.¹,² Because a rejuvenation therapy is assembled from several independently patentable pieces, the reprogramming-factor set and its ratios, the delivery system, the inducible control mechanism, the target tissue and indication, and the tools used to measure biological age, freedom-to-operate is a multi-layer, multi-owner analysis rather than a single clearance.
The foundational layer shapes everything above it. The original induced-pluripotent-stem-cell reprogramming methods, established through the forced expression of a defined set of transcription factors, sit under a well-known foundational estate that has been broadly licensed,³ and partial-reprogramming approaches inherit questions about how far that foundation reaches. Independent work has shown that epigenetic reprogramming can unlock tissue regenerative potential, reinforcing why these methods are so contested.⁴ This academic origin is visible in the ownership record: across the Cypris corpus of more than 500 million patents and scientific papers, the most active assignees in the cellular-reprogramming and induced-pluripotency space are led by academic and translational institutions, including Kyoto University, the University of California San Diego, the University of Texas System, Memorial Sloan Kettering, and Harvard, alongside cell-therapy companies, and the corpus holds on the order of 28,700 de-duplicated families, with the United States, China, and Japan the leading jurisdictions. Layered on top are newer, fast-growing estates specific to partial and transient reprogramming, cyclic and inducible expression schemes, chemical or small-molecule reprogramming that avoids transcription factors altogether, and tissue-specific delivery. Because applications publish about eighteen months after filing, the most recent reprogramming, delivery, and control filings are under-represented, so the current frontier is more active than granted-patent counts suggest.
The landscape is a well-capitalized race, and the strategic question is which layer to own. In January 2026 the field reached a milestone when the US Food and Drug Administration cleared the first human trial of a partial epigenetic reprogramming therapy, an investigational optic-neuropathy treatment; the clearance authorizes a first-in-human study and is not itself evidence of efficacy.⁹ Across the Cypris corpus, filings in this space grew from a few hundred families per year at the start of the last decade to roughly 3,200 in 2024, with 2025 counts partial because of the publication lag. Several richly funded companies are pursuing different factor sets, delivery routes, and target tissues, and a recurring challenge is to separate genuine rejuvenation, a measured reduction in biological age, from a mere slowing of decline.⁵ The durable value increasingly sits not in the general idea of reprogramming, which rests on the contested foundation, but in the specific, well-supported improvements: safe and controllable expression systems that avoid tumor risk, factor combinations and chemical alternatives, tissue-targeted delivery, and the validated biomarkers, including epigenetic clocks, used to demonstrate rejuvenation.⁶,⁷,⁸ Reading the landscape by layer and by owner, and tracking both the patents and the underlying research, is what separates a workable position from a blocked one.
What creates FTO risk in cellular reprogramming
Foundational reprogramming claims. These cover the underlying induced-pluripotency methods and factor sets, a broadly licensed foundation whose reach into partial approaches shapes everything above it.
Partial and inducible-control claims. These cover transient, cyclic, and inducible expression schemes that rejuvenate without full dedifferentiation, a fast-growing and contested layer.
Delivery claims. These cover viral vectors, lipid nanoparticles, and mRNA delivery of reprogramming factors, a distinct and separately owned layer often decisive for a therapy.
Chemical and small-molecule reprogramming claims. These cover approaches that induce rejuvenation without transcription factors, an emerging and less-crowded route.
Target, indication, and biomarker claims. These cover specific tissues and indications and the epigenetic-age measurements used to demonstrate effect, so a platform can be free for one application and blocked for another.
How AI-powered landscape and FTO analysis helps
A multi-layer, multi-owner landscape built on a contested foundation is beyond manual clearance. AI-powered analysis addresses this with semantic search that retrieves relevant foundational, partial-reprogramming, delivery, control, and target claims regardless of terminology, attribution that resolves academic and commercial owners to canonical entities and captures the license and spinout chains, claim-level analysis that separates the layers, and continuous monitoring that tracks new filings and the fast-moving research. Because reprogramming advances appear in scientific literature well before they are patented, reading both patents and literature gives the earliest warning of where the field is heading.
Where Cypris fits
Cypris runs patent landscape and freedom-to-operate analysis for multi-layer, academically rooted fields such as cellular reprogramming across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters the landscape by layer, foundational reprogramming, partial and inducible control, delivery, chemical reprogramming, and target and biomarker, and normalizes academic and commercial owners to canonical entities, so a team can trace how rights and licenses are distributed across many parties rather than read a flat list. Semantic search across patents and scientific literature surfaces relevant claims regardless of terminology and connects filings to the underlying research, which is where new factor sets, control systems, and delivery methods emerge first. Cypris Q, the platform's agentic layer, lets teams run landscape and FTO analysis conversationally and chain the attribution, clustering, and claim-level analysis across layers, and Agentic Monitoring tracks the landscape over time and flags new filings and developments as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is cellular reprogramming in the longevity context? Cellular reprogramming in the longevity context is the use of reprogramming factors to reset the epigenetic state of aged cells toward a younger profile. Partial or transient reprogramming applies a subset of the Yamanaka factors briefly, aiming to rejuvenate cells without erasing their identity. It is being pursued as an approach to age-related disease and tissue restoration.
Why is freedom-to-operate hard for reprogramming therapies? Freedom-to-operate is hard for reprogramming therapies because a therapy is assembled from several independently patentable layers, the reprogramming-factor set, the delivery system, the inducible control mechanism, the target tissue, and biomarker tools, often held by different owners on top of a foundational estate. Clearing one layer does not clear the others. FTO is therefore a multi-layer, multi-owner analysis.
How does the foundational iPSC estate affect partial reprogramming? The foundational induced-pluripotent-stem-cell estate affects partial reprogramming because partial approaches use the same reprogramming factors, so questions about how far the foundation reaches propagate into the newer methods. The foundation has been broadly licensed. Partial-reprogramming developers must consider both the foundation and the specific improvement layers.
What claim types create FTO risk in reprogramming? Five claim types create FTO risk: foundational reprogramming claims, partial and inducible-control claims, delivery claims, chemical and small-molecule reprogramming claims, and target, indication, and biomarker claims. Each covers a distinct layer and can be held by a different owner. Control systems and delivery are especially decisive.
Where is the white space in cellular reprogramming? The white space sits in safe and controllable expression systems that avoid tumor risk, chemical and small-molecule reprogramming, tissue-specific delivery, specific factor combinations, and validated biomarkers of biological age. The general concept rests on a contested foundation. The durable, defensible value is in these specific improvement and delivery layers.
Why does reprogramming analysis need scientific literature? Reprogramming analysis needs scientific literature because new factor sets, control systems, and delivery methods appear in research well before they are patented, so the literature gives the earliest signal in a fast-moving field. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
What software helps analyze the cellular reprogramming patent landscape? Software for the cellular reprogramming landscape should resolve academic and commercial owners and license chains to canonical entities, cluster the foundational, control, delivery, and target layers, search patents and scientific literature semantically, and monitor a fast-moving field continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams need reprogramming patent landscape and FTO analysis? Reprogramming patent landscape and FTO analysis is needed by R&D, IP, and business-development teams at longevity and gene-therapy companies, academic technology-transfer offices, and investors assessing rejuvenation assets. The multi-layer, contested landscape makes structured analysis essential. Cypris serves hundreds of enterprise customers across pharmaceuticals and other research-intensive industries.
This article addresses patents and freedom-to-operate and is not legal, medical, or investment advice, and contains no clinical or dosing guidance. FTO determinations should be reviewed with qualified patent counsel.
Endnotes
- Ocampo, A., Reddy, P., Izpisua Belmonte, J. C., et al. (2016). In vivo amelioration of age-associated hallmarks by partial reprogramming. Cell, 167(7). https://doi.org/10.1016/j.cell.2016.11.052
- Lu, Y., Krishnan, A., Sinclair, D. A., et al. (2020). Reprogramming to recover youthful epigenetic information and restore vision. Nature, 588. https://doi.org/10.1038/s41586-020-2975-4
- Takahashi, K., & Yamanaka, S. (2013). Induced pluripotent stem cells in medicine and biology. Development, 140(12). https://doi.org/10.1242/dev.092551
- Reddy, P., Izpisua Belmonte, J. C., & Memczak, S. (2021). Unlocking tissue regenerative potential by epigenetic reprogramming. Cell Stem Cell, 28(3). https://doi.org/10.1016/j.stem.2020.12.006
- Zhang, B., Trapp, A., Kerepesi, C., & Gladyshev, V. N. (2021). Emerging rejuvenation strategies—reducing the biological age. Aging Cell, 21(1). https://doi.org/10.1111/acel.13538
- Moqri, M., Poganik, J. R., Gladyshev, V. N., & Horvath, S. (2025). What makes biological age epigenetic clocks tick. Nature Aging. https://doi.org/10.1038/s43587-025-00833-1
- Mammalian Methylation Consortium; Horvath, S., et al. (2023). Universal DNA methylation age across mammalian tissues. Nature Aging, 3. https://doi.org/10.1038/s43587-023-00462-6
- Ferrucci, L., et al. (2019). Measuring biological aging in humans: a quest. Aging Cell, 19(2). https://doi.org/10.1111/acel.13080
- Life Biosciences (2026, January 28). Life Biosciences announces FDA clearance of IND application for ER-100 in optic neuropathies. https://www.lifebiosciences.com/life-biosciences-announces-fda-clearance-of-ind-application-for-er-100-in-optic-neuropathies

Prior art search for artificial intelligence and machine learning inventions is one of the hardest retrieval problems in patent work, for reasons specific to how AI knowledge is produced and disclosed. Prior art search establishes whether an invention is novel by finding any earlier disclosure that describes it. In most fields, the relevant disclosures are predominantly patents. In AI and machine learning, the most relevant and most recent disclosures are predominantly non-patent literature: preprints on arXiv, proceedings from conferences such as NeurIPS and ICML, open-source code and model documentation, and technical reports. These sources are published quickly and openly, often well ahead of any corresponding patent, so a prior art search confined to patent databases misses the state of the art.
The volume compounds the difficulty. AI scientific publications more than doubled from about 102,000 in 2013 to more than 242,000 in 2023, growing nearly 20 percent in the final year alone.¹ Patenting has grown even faster from a smaller base: AI patents granted worldwide rose from 3,833 in 2010 to 122,511 in 2023, an increase of almost 30 percent in the last year measured.¹ Generative AI illustrates the velocity of the literature most sharply, with related scientific publications rising from 116 in 2014 to more than 34,000 in 2023 while generative-AI patent families grew more than 800 percent over roughly the same period.² A prior art searcher in this field is therefore working against both a large and a rapidly expanding corpus, split across patent and non-patent sources.
Retrieval quality falls exactly where AI prior art needs it most. Patent retrieval is already harder than general-domain information retrieval, and controlled evaluation shows that cross-domain retrieval, finding relevant art outside the query's own technology area, performs several times worse than in-domain retrieval; one recent family-level benchmark found out-of-domain retrieval roughly five times worse than in-domain across hundreds of controlled configurations.³,⁴ AI and machine-learning methods are applied across many application domains, so relevant prior art for an AI invention is frequently located in a different field than the invention's stated use, which is precisely the cross-domain case where conventional retrieval degrades. This is the technical reason keyword and classification search alone are insufficient for AI prior art, and why dense, semantic methods have become the focus of research on patent prior art retrieval.⁵,⁶
Why AI prior art is distinctively hard
Non-patent literature dominates. The most relevant and most recent AI disclosures appear first in preprints, conference proceedings, and open-source code, so a patent-only search misses the state of the art.
Exploding volume. AI publications more than doubled to over 242,000 in 2023, and AI patents granted rose to 122,511, so the corpus a searcher must cover is both large and expanding rapidly.¹
Cross-domain dispersion. AI methods are applied across many fields, so relevant prior art is often in a different technology area than the invention, which is where retrieval degrades most.³
Fast obsolescence of terminology. AI vocabulary evolves quickly, so keyword search misses conceptually identical work described in newer or different terms.
Software-claim breadth. Algorithmic and software claims can be drafted broadly and abstractly, which makes matching a claim to its closest prior art a conceptual rather than a lexical task.
How semantic search closes the gap
Semantic search addresses each of these problems. It retrieves conceptually relevant disclosures regardless of terminology, which handles both fast-evolving vocabulary and broadly drafted software claims. Applied across both patents and scientific literature in one corpus, it covers the non-patent literature where AI prior art concentrates rather than patents alone. And because dense retrieval encodes meaning rather than surface form, it is better positioned than keyword search for the cross-domain case, retrieving relevant art from a different application area than the invention. Combined with an ontology that organizes retrieval by concept, semantic search returns a structured, high-recall view of the prior art rather than a keyword-limited sample.
Where Cypris fits
Cypris runs semantic prior art search across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Because the corpus spans both patents and scientific literature, Cypris covers the non-patent literature where AI and machine-learning prior art concentrates, rather than patents alone. Semantic search retrieves conceptually relevant disclosures regardless of terminology, which handles the fast-evolving vocabulary and broadly drafted software claims characteristic of AI inventions, and the ontology organizes retrieval by concept so cross-domain prior art in a different application area is surfaced rather than missed. Cypris Q, the platform's agentic layer, lets teams run and chain prior art and novelty analysis conversationally, and Agentic Monitoring tracks a technology area over time so newly published disclosures are surfaced as they appear, which matters in a field moving as fast as AI. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is prior art search hard for AI and machine learning inventions?
Prior art search is hard for AI and machine learning inventions because the most relevant and most recent disclosures are predominantly non-patent literature, such as preprints, conference proceedings, and open-source code, which a patent-only search misses. The corpus is also large and expanding rapidly, and AI methods are dispersed across many application domains. These factors make high-recall, cross-domain retrieval essential.
Why does non-patent literature matter so much for AI prior art?
Non-patent literature matters for AI prior art because AI research is published quickly and openly, often well ahead of any corresponding patent, so the state of the art appears first in preprints, conference papers, and code. A search confined to patent databases misses these disclosures. Effective AI prior art search must cover both patents and scientific literature.
How large is the AI prior art corpus?
The AI prior art corpus is large and growing quickly. AI scientific publications more than doubled from about 102,000 in 2013 to over 242,000 in 2023, and AI patents granted worldwide rose from 3,833 in 2010 to 122,511 in 2023. Generative-AI publications alone grew from 116 in 2014 to more than 34,000 in 2023.
What makes AI prior art retrieval technically difficult?
AI prior art retrieval is technically difficult because AI methods are applied across many domains, so relevant prior art is often in a different technology area than the invention, and cross-domain retrieval performs several times worse than in-domain retrieval. One benchmark found out-of-domain retrieval roughly five times worse than in-domain. Fast-evolving terminology and broadly drafted software claims add further difficulty.
Why is keyword search insufficient for AI prior art?
Keyword search is insufficient for AI prior art because AI terminology evolves quickly and software claims are often drafted broadly and abstractly, so conceptually identical work is described in different terms. Keyword search matches surface form and misses these. Semantic search retrieves by meaning, which is what the task requires.
How does semantic search improve AI prior art search?
Semantic search improves AI prior art search by retrieving conceptually relevant disclosures regardless of terminology, across both patents and scientific literature, and by handling the cross-domain case where relevant art is in a different field. It encodes meaning rather than surface form. Combined with an ontology, it returns a structured, high-recall view of the prior art.
Does AI prior art search need to cover scientific literature?
AI prior art search needs to cover scientific literature because the most relevant and most recent AI disclosures appear there first, in preprints, conference proceedings, and technical reports. Covering patents alone leaves the state of the art unretrieved. Cypris searches both across more than 500 million patents and scientific papers.
Which teams run AI prior art search?
AI prior art search is run by IP, R&D, and patent teams at technology companies and across industries adopting AI, as well as by patent professionals assessing novelty. It is increasingly important as AI patenting grows. Cypris serves hundreds of enterprise customers across research-intensive and regulated industries.
How current does AI prior art search need to be?
AI prior art search needs to be continuously current, because AI research and filings publish constantly and the state of the art shifts quickly. A one-time search reflects only the moment it was run. Cypris uses Agentic Monitoring to track a technology area and surface newly published disclosures as they appear.
Endnotes
- Stanford Institute for Human-Centered Artificial Intelligence (2025). Artificial Intelligence Index Report 2025, Chapter 1. arXiv:2504.07139. https://doi.org/10.48550/arxiv.2504.07139
- World Intellectual Property Organization (2024). Patent Landscape Report: Generative Artificial Intelligence. Geneva: WIPO. https://doi.org/10.34667/tind.49740
- Cavallucci, N., Chibane, I. & Ayaou, M. (2026). DAPFAM: A Domain-Aware Family-level Dataset to benchmark cross-domain patent retrieval. Array. https://doi.org/10.1016/j.array.2026.100720
- Lupu, M. (2013). Patent Retrieval. Foundations and Trends in Information Retrieval. https://doi.org/10.1561/1500000027
- Stamatis, V. (2022). End to End Neural Retrieval for Patent Prior Art Search. Lecture Notes in Computer Science. https://doi.org/10.1007/978-3-030-99739-7_66
- Zihayat, M. & Etwaroo, R. (2021). A non-factoid question answering system for prior art search. Expert Systems with Applications. https://doi.org/10.1016/j.eswa.2021.114910
