A faster, more accurate way to explore innovation data—now available in Cypris.
For innovation teams, speed and accuracy aren’t optional—they’re critical. You need to quickly find all relevant documents, slice and dice datasets however you want, and trust that the results are complete and representative. With this in mind, we’ve upgraded how semantic search works inside Cypris.
Today, we’re launching an upgraded search infrastructure that gives users access to full, exact result sets—unlocking more powerful analysis, faster iteration, and deterministic filtering and charting.
Unlike traditional semantic or vector search engines—which make it difficult to count, filter, or chart large sets of matched documents—our new approach prioritizes transparency and performance while preserving semantic relevance.
Why we moved away from vector search
Our original implementation relied on semantic and vector search to capture the “meaning” behind user queries. But as our platform evolved, it became clear that these systems weren’t well-suited for our core use cases.
Users needed:
- Deterministic filtering (e.g., "how many results match this atom?")
- Transparent, complete result sets to power charts and dashboards
- Fast, repeatable queries that don’t change subtly over time
Modern vector search systems don’t easily support this level of transparency. They return approximate matches and abstract similarity scores, often making it hard to understand why a document was returned—or whether it’s the full picture.
So we made a decision: move away from vector search and lean into what traditional search engines do best.
A return to boolean and lexical search—with a twist
We rebuilt our search infrastructure on top of Elasticsearch’s powerful boolean and lexical search capabilities. This shift brings major advantages:
- Faster query speeds that dramatically improve iteration time
- Deterministic filtering and counts, so every chart is grounded in the full dataset
- Predictable, explainable results that users can trust
But we didn’t stop there.
To preserve the benefits of semantic understanding, we’ve rethought where that intelligence should live—not at query time, but at data ingestion.
Capturing semantic meaning at ingest time
Instead of computing document-query similarity during search, we enrich documents at the time of ingestion. Here’s how:
- Synonym expansion: We find related words and concepts not explicitly mentioned in the document and add them as fields, enabling semantic-style recall via lexical search.
- Stemming: Both queries and documents are reduced to their root forms, allowing consistent matches (e.g., “running” and “run”).
The result? You get the same functionality—semantically relevant results—without the opacity or latency tradeoffs of vector search.
What’s next: Reranking for even better relevance
We’re not done. Coming soon to Cypris is a reranking layer that boosts the most relevant results to the top of the list using lightweight vector techniques.
Here’s how it works:
- A standard lexical search retrieves the full result set.
- We take the top N results and rerank them using vector similarity, powered by Elasticsearch’s new hybrid scoring capabilities.
- You get faster queries with even better relevance—without compromising on counts or transparency.
This layered approach gives us the best of both worlds: precise filtering and fast queries, plus smarter ordering of results where it matters most.
We’re excited to bring this upgrade to our users, and we’re already seeing teams iterate faster and uncover insights more confidently. This is a foundational shift—and just the beginning of what’s to come.
Want a walkthrough of what’s changed? Reach out to our team.

Introducing our upgraded semantic search
A faster, more accurate way to explore innovation data—now available in Cypris.
For innovation teams, speed and accuracy aren’t optional—they’re critical. You need to quickly find all relevant documents, slice and dice datasets however you want, and trust that the results are complete and representative. With this in mind, we’ve upgraded how semantic search works inside Cypris.
Today, we’re launching an upgraded search infrastructure that gives users access to full, exact result sets—unlocking more powerful analysis, faster iteration, and deterministic filtering and charting.
Unlike traditional semantic or vector search engines—which make it difficult to count, filter, or chart large sets of matched documents—our new approach prioritizes transparency and performance while preserving semantic relevance.
Why we moved away from vector search
Our original implementation relied on semantic and vector search to capture the “meaning” behind user queries. But as our platform evolved, it became clear that these systems weren’t well-suited for our core use cases.
Users needed:
- Deterministic filtering (e.g., "how many results match this atom?")
- Transparent, complete result sets to power charts and dashboards
- Fast, repeatable queries that don’t change subtly over time
Modern vector search systems don’t easily support this level of transparency. They return approximate matches and abstract similarity scores, often making it hard to understand why a document was returned—or whether it’s the full picture.
So we made a decision: move away from vector search and lean into what traditional search engines do best.
A return to boolean and lexical search—with a twist
We rebuilt our search infrastructure on top of Elasticsearch’s powerful boolean and lexical search capabilities. This shift brings major advantages:
- Faster query speeds that dramatically improve iteration time
- Deterministic filtering and counts, so every chart is grounded in the full dataset
- Predictable, explainable results that users can trust
But we didn’t stop there.
To preserve the benefits of semantic understanding, we’ve rethought where that intelligence should live—not at query time, but at data ingestion.
Capturing semantic meaning at ingest time
Instead of computing document-query similarity during search, we enrich documents at the time of ingestion. Here’s how:
- Synonym expansion: We find related words and concepts not explicitly mentioned in the document and add them as fields, enabling semantic-style recall via lexical search.
- Stemming: Both queries and documents are reduced to their root forms, allowing consistent matches (e.g., “running” and “run”).
The result? You get the same functionality—semantically relevant results—without the opacity or latency tradeoffs of vector search.
What’s next: Reranking for even better relevance
We’re not done. Coming soon to Cypris is a reranking layer that boosts the most relevant results to the top of the list using lightweight vector techniques.
Here’s how it works:
- A standard lexical search retrieves the full result set.
- We take the top N results and rerank them using vector similarity, powered by Elasticsearch’s new hybrid scoring capabilities.
- You get faster queries with even better relevance—without compromising on counts or transparency.
This layered approach gives us the best of both worlds: precise filtering and fast queries, plus smarter ordering of results where it matters most.
We’re excited to bring this upgrade to our users, and we’re already seeing teams iterate faster and uncover insights more confidently. This is a foundational shift—and just the beginning of what’s to come.
Want a walkthrough of what’s changed? Reach out to our team.

Keep Reading

Green steel has become one of the most closely watched areas of industrial decarbonization, and its patent landscape is distinctive because low-carbon steelmaking is not a single technology but a set of competing routes, each with its own chemistry and process engineering. Conventional steelmaking reduces iron ore with coal-derived coke in a blast furnace, and ironmaking generates roughly 7 percent of global CO2 emissions across an industry producing about 1.85 billion tonnes of steel a year<sup>2</sup>. The leading low-carbon routes replace that chemistry in different ways, and each is a distinct region of patenting: hydrogen-based direct reduction uses green hydrogen instead of coke to turn iron ore into sponge iron, which is then melted in an electric arc furnace<sup>3</sup>; molten oxide electrolysis passes electricity through molten iron ore, producing liquid metal and oxygen at the anode with no process CO2 given a clean electricity input<sup>4</sup>; and low-temperature electrochemical routes produce iron from ore or low-grade feedstocks by electrowinning, though the aqueous chemistry still faces a hydrogen-evolution-reaction efficiency bottleneck that limits faradaic efficiency<sup>7</sup>. Because each route relies on different core steps, anode and electrolyte materials, hydrogen integration, ore handling, and furnace design, freedom-to-operate and white space analysis must treat green steel as several landscapes at once.
The field is moving from pilots to first industrial-scale plants. A hydrogen direct-reduction plant designed for a developer-reported emissions reduction of up to roughly 95 percent versus blast-furnace production — a figure consistent with, though not itself drawn from, peer-reviewed techno-economic modeling of the H2-DRI/EAF route<sup>1</sup> — is being built at industrial scale and is on track to begin production, and electrolysis-based developers are scaling reactors toward commercial output. The intellectual property reflects the maturity gap between the routes: hydrogen direct reduction builds on established direct-reduced-iron practice and concentrates IP in hydrogen integration, reduction control, and furnace operation, with break-even hydrogen pricing as a central techno-economic question in the peer-reviewed literature<sup>1</sup>, while the electrolysis routes concentrate foundational IP in the inert-anode and electrolyte materials and cell designs that make emission-free iron production work<sup>4,5</sup>, much of it traceable to a small number of academic and company lineages. Because applications publish about eighteen months after filing, the most recent electrolysis and process filings are under-represented, so the current frontier is more active than granted-patent counts suggest.
The strategic question is which route and layer to back, and the white space sits where cost, materials, and feedstock constraints are hardest. In hydrogen direct reduction, the open ground is in reducing hydrogen consumption and cost, tolerating lower-grade ore, and integrating variable hydrogen supply<sup>1</sup>. In molten oxide electrolysis, durable inert-anode materials that survive the process are the central, high-value problem<sup>4</sup>. In low-temperature electrowinning, the opportunity is in efficient electrochemistry and the use of low-grade ores and mining waste, with comparative techno-economic analysis showing how the three electrolysis-adjacent routes trade off against hydrogen reduction<sup>6,7</sup>. Across all routes, ore flexibility is strategically important because some routes require scarce high-grade ore. Reading the landscape by route, core step, and owner, and tracking both the patents and the underlying process research, is what separates a crowded region from an open one.
Where the green-steel white space is
Inert-anode and electrolyte materials. Durable anode and electrolyte materials that survive molten oxide electrolysis are the central, high-value problem for the electrolysis route<sup>4,5</sup>.
Low-grade ore tolerance. Processes that use lower-grade ore or mining waste ease the feedstock constraint that limits some routes and broaden where plants can be sited.
Hydrogen integration and reduction control. Reducing hydrogen consumption and cost and integrating variable green-hydrogen supply in direct reduction is a large, active layer, with break-even hydrogen price as the key economic lever<sup>1</sup>.
Low-temperature electrochemical iron production. Efficient aqueous-phase electrowinning of iron is an earlier, less-crowded route with distinct chemistry, currently constrained by hydrogen-evolution-reaction efficiency losses<sup>7</sup>.
Furnace and process integration. Integrating direct-reduced iron with electric arc furnaces and optimizing continuous operation is where cost and quality are decided<sup>3</sup>.
How AI-powered landscape and white space analysis helps
Resolving a landscape that spans several production routes, each with its own chemistry and process, requires more than keyword search. AI-powered analysis addresses this with semantic search that clusters activity by route, core step, and material across varied terminology, attribution that normalizes filers to canonical entities and tracks new entrants, and continuous monitoring that keeps pace with a fast-commercializing field. Because green-steel advances appear in scientific and process-engineering literature before they are patented, reading both patents and literature gives the earliest signal of where scalable routes are emerging.
The competitive landscape by the numbers
Cypris's corpus puts the low-carbon steelmaking patent family set — spanning hydrogen-DRI, electrolysis/molten oxide electrolysis, electrowinning, and general "green steel" filings — at roughly 27,049 families (Cypris corpus, indicative; 2025–26 partial). Filing has run at roughly 900–1,900 new families per year across 2016–2024, with 2025 (2,210) and 2026 (1,750, partial) continuing the trend (Cypris corpus, indicative; 2025–26 partial). The assignee ranking spans both steel majors and petrochemical/catalysis houses: Sinopec (431 families), Nippon Steel (229), ArcelorMittal (215), JFE (98), and Northeastern University (111) lead the count (Cypris corpus, indicative; 2025–26 partial) — worth flagging, since several of the top filers are catalysis and process-engineering companies rather than primary steelmakers, so the set is broader than steel production alone. Geographically, China dominates with 14,065 families, followed by the United States (1,233), Germany (636), Japan (329), Luxembourg (271, reflecting ArcelorMittal's filings), and Sweden (151) (Cypris corpus, indicative; 2025–26 partial).
Where Cypris fits
Cypris runs patent landscape and white space analysis for multi-route industrial fields such as green steel across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by route, hydrogen direct reduction, molten oxide electrolysis, and electrowinning, and by layer, anode and electrolyte, hydrogen integration, ore handling, and furnace design, and normalizes filers to canonical entities, so a team can resolve which routes and layers are crowded and which remain open as white space, and can track new entrants as the field scales. Semantic search across patents and scientific literature connects filings to the underlying process and materials research, which is where green-steel advances appear first. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined route over time and flags new patents and papers as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is the green steel patent landscape? The green steel patent landscape is the set of patents covering low-carbon steelmaking. It divides across competing routes, hydrogen-based direct reduction feeding an electric arc furnace, molten oxide electrolysis, and low-temperature electrowinning, each with distinct chemistry and process IP<sup>1,4,7</sup>. Each route is a distinct region of patenting.
Why is steelmaking a decarbonization priority? Steelmaking is a decarbonization priority because ironmaking generates roughly 7 percent of global CO2 emissions across an industry producing about 1.85 billion tonnes of steel a year<sup>2</sup>. Low-carbon routes replace coke-based reduction with hydrogen or electricity. The first industrial-scale plants are now being built.
What routes does the green-steel landscape cover? The landscape covers hydrogen-based direct reduced iron, which uses green hydrogen instead of coke<sup>3</sup>; molten oxide electrolysis, which splits molten iron ore with electricity to yield liquid metal and oxygen<sup>4</sup>; and low-temperature electrochemical iron production by electrowinning<sup>6,7</sup>. Each relies on different core steps and materials. Freedom-to-operate and white space analysis must treat them separately.
Where is the white space in green steel? The white space includes inert-anode and electrolyte materials for electrolysis, low-grade ore tolerance, hydrogen integration and reduction control, low-temperature electrochemical iron production, and furnace and process integration. The routes sit at different maturity levels. The most open, high-value opportunities are in the electrolysis materials and in ore and hydrogen flexibility.
Why are inert-anode materials so important? Inert-anode materials are important because molten oxide electrolysis depends on an anode that can survive extreme temperatures and produce oxygen rather than carbon dioxide, and finding durable, affordable anode and electrolyte materials is the central technical problem for that route<sup>4,5</sup>. Solving it is what makes emission-free electrolytic iron viable. Much of the route's defensible IP concentrates there.
Is molten oxide electrolysis actually "zero-carbon"? Molten oxide electrolysis is more precisely described as producing oxygen and liquid metal with no process CO2, provided the electricity input is clean — the process itself does not emit carbon during reduction, but the claim depends on the power source<sup>4</sup>. Unqualified "zero-carbon" framing overstates this without specifying the electricity mix. That distinction matters for both technical and disclosure purposes.
Who is filing green-steel patents, and where? In Cypris's corpus of roughly 27,049 low-carbon steelmaking patent families, China dominates filing activity, followed by the United States, Germany, Japan, and Luxembourg, and the assignee ranking includes both steel majors (Nippon Steel, ArcelorMittal, JFE) and petrochemical/catalysis filers (Sinopec) (Cypris corpus, indicative; 2025–26 partial).
Why does green-steel analysis need scientific literature? Green-steel analysis needs scientific literature because reduction, electrolysis, and materials advances appear in process research before they are patented, so the literature gives the earliest signal. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
What software helps analyze the green steel patent landscape? Software for the green-steel landscape should cluster activity by route and process layer, resolve filers to canonical owners, search patents and scientific literature semantically, and monitor a fast-commercializing field continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams use green steel patent landscape analysis? Green steel patent landscape analysis is used by R&D, innovation, IP, and strategy teams at steelmakers, mining and materials companies, electrolysis and hydrogen developers, and their partners, as well as investors and policymakers. It informs which route to back, where to file, and where competitors are concentrated. Cypris serves hundreds of enterprise customers across advanced materials, energy, chemicals, and other regulated industries.
Endnotes
- Papadias DD, Brooks K, Yoro KO, Autrey T, et al. (Argonne National Laboratory, Lawrence Berkeley National Laboratory, Pacific Northwest National Laboratory; DOE-funded). Green steel: design and cost analysis of hydrogen-based direct iron reduction. Energy & Environmental Science. 2023. DOI: 10.1039/d3ee01077e.
- Bae JW, Raabe D, et al. Reducing iron oxide with ammonia: a sustainable path to green steel. Advanced Science. 2023. DOI: 10.1002/advs.202300111.
- Boretti A. The perspective of hydrogen direct reduction of iron. Journal of Cleaner Production. 2023. DOI: 10.1016/j.jclepro.2023.139585.
- Paramore JD, Kim H, Allanore A, Sadoway DR (MIT). Stability of iridium anode in molten oxide electrolysis for ironmaking. ECS Transactions. 2010. DOI: 10.1149/1.3484779.
- Azimi G, Allanore A, Judge WD, Sadoway DR. E-logpO2 diagrams for ironmaking by molten oxide electrolysis. Electrochimica Acta. 2017. DOI: 10.1016/j.electacta.2017.07.059.
- Rhamdhani MA, et al. (CSIRO, Swinburne University). Economics of electrowinning iron from ore for green steel production. Journal of Sustainable Metallurgy. 2024. DOI: 10.1007/s40831-024-00878-3.
- Viswanathan V, Kavalsky L. Electrowinning for room-temperature ironmaking: mapping the electrochemical aqueous iron interface. Journal of Physical Chemistry C. 2024. DOI: 10.1021/acs.jpcc.4c01867.
- Cypris platform corpus analysis, low-carbon steelmaking patent families. Indicative figures; 2025–2026 partial.

Targeted protein degradation has become one of the most closely watched modalities in drug discovery, and its patent landscape is distinctive because a degrader is a modular molecule whose parts are patented separately. Rather than blocking a protein's active site the way a conventional inhibitor does, a degrader recruits the cell's ubiquitin-proteasome system to destroy the target protein outright, which makes it possible to address targets that lack a druggable pocket.¹ The two most advanced approaches are proteolysis-targeting chimeras, or PROTACs, which are heterobifunctional molecules built from a ligand that binds the target protein, a linker, and a ligand that binds an E3 ubiquitin ligase, and molecular glues, which are smaller, single-piece molecules that induce proximity between the target and an E3 ligase by reprogramming the ligase's surface to recruit a neosubstrate.²,³ A growing set of related modalities, including lysosome-targeting and autophagy-targeting chimeras and degrader-antibody conjugates, extends the field further, and the chemical space of molecular glues in particular is only beginning to be mapped.⁴ Because the E3-ligase binder, the target ligand, the linker, and the whole composite molecule can each be claimed independently and are often held by different owners, freedom-to-operate for a degrader is a multi-layer, multi-owner analysis rather than a single clearance.
The field has moved from concept to the market, which has raised the stakes across every layer. In May 2026 vepdegestrant (VEPPANU), an oral PROTAC estrogen-receptor degrader developed by Arvinas and Pfizer, received US Food and Drug Administration approval, becoming the first approved PROTAC therapy, and the partners had earlier moved to out-license its commercialization.⁸,⁹ A steady stream of degrader deals has followed, including a second Monte Rosa–Novartis molecular-glue collaboration announced in September 2025 with a $120 million upfront payment and total potential value up to $5.7 billion.¹⁰ Foundational intellectual property traces to the academic origins of the PROTAC concept and to the E3-ligase-binder chemistries: the large majority of clinical-stage degraders recruit the cereblon ligase, with von Hippel-Lindau the other principal handle, even as the field works to expand to the other canonical E3 ligases and beyond.³,⁵ Patent activity around the von Hippel-Lindau layer alone is now substantial enough to sustain dedicated patent reviews.⁷ This concentration is visible in the record: across the Cypris corpus of more than 500 million patents and scientific papers, the degrader set holds on the order of 4,292 families and grew from roughly 134 in 2020 to about 814 in 2024, with the most active assignees including Dana-Farber, C4 Therapeutics, and Arvinas, and China (about 1,492 families) modestly ahead of the United States (about 1,205); 2025 and 2026 counts are partial because of the publication lag.
The strategic picture turns on where defensible, hard-to-design-around IP sits. The human genome encodes more than six hundred E3 ligases, but only a handful have been harnessed for degradation, so novel E3-ligase binders are a high-value, comparatively open layer, and the molecular-glue field, where rational design is still early, is another.¹,⁶ Because a degrader assembled from a known target ligand and a known E3 binder may face freedom-to-operate exposure on either component plus the linker, the durable value increasingly lies in new E3 chemistries, glue scaffolds, tissue- or ligase-selective designs, orally bioavailable degraders, and expansion beyond oncology into immunology and neuroscience.² Reading the landscape by layer and by owner, and tracking both the patents and the underlying chemistry and cell-biology research, is what separates a workable position from a blocked one.
What creates FTO risk in targeted protein degradation
E3-ligase-binder claims. These cover the chemistries that recruit an E3 ligase, such as cereblon and von Hippel-Lindau binders and newer ligases, a foundational and heavily contested layer.³,⁷
Target-ligand claims. These cover the warhead that binds the protein of interest, which can carry its own separate IP from inhibitor programs.
Linker claims. These cover the chemistry connecting the two ligands in a PROTAC, a distinct layer that materially affects degradation and is independently patentable.
Composite-molecule and molecular-glue claims. These cover the specific bifunctional degrader or single-piece glue, the layer most directly tied to a clinical candidate.²
Mechanism, formulation, and modality claims. These cover degradation mechanisms, formulations, and emerging modalities such as lysosome-targeting chimeras and degrader-antibody conjugates.
How AI-powered landscape and FTO analysis helps
A modular, multi-owner, fast-moving landscape is beyond manual clearance. AI-powered analysis addresses this with semantic search that retrieves relevant E3-binder, target-ligand, linker, and composite-molecule claims regardless of terminology, attribution that resolves academic and commercial owners and the license chains to canonical entities, claim-level analysis that separates the layers, and continuous monitoring that tracks new filings and deals. Because degrader advances appear in scientific literature before they are patented, reading both patents and literature gives earlier warning of where the field is heading.
Where Cypris fits
Cypris runs patent landscape and freedom-to-operate analysis for modular, contested fields such as targeted protein degradation across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters the landscape by layer, E3-ligase binder, target ligand, linker, and composite molecule, and normalizes academic and commercial owners to canonical entities, so a team sees how rights are distributed across the many parties rather than a flat list. Semantic search across patents and scientific literature surfaces relevant claims regardless of terminology and connects filings to the underlying research, which is where new E3 chemistries and glue scaffolds emerge first. Cypris Q, the platform's agentic layer, lets teams run landscape and FTO analysis conversationally and chain the attribution, clustering, and claim-level analysis across layers, and Agentic Monitoring tracks the landscape over time and flags new filings and developments as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is targeted protein degradation? Targeted protein degradation is a modality that eliminates a disease-causing protein by recruiting the cell's ubiquitin-proteasome system, rather than inhibiting the protein's activity. The leading approaches are PROTACs, which are bifunctional molecules, and molecular glues, which are single-piece molecules. It can address targets that lack a druggable pocket.
Why is freedom-to-operate hard for degraders? Freedom-to-operate is hard for degraders because a PROTAC is built from an E3-ligase binder, a target ligand, and a linker, each independently patentable and often held by different owners, and the composite molecule is a further layer. Molecular glues add their own scaffold IP. FTO must therefore be assessed layer by layer across multiple estates.
What claim types create FTO risk in TPD? Five claim types create FTO risk? E3-ligase-binder claims, target-ligand claims, linker claims, composite-molecule and molecular-glue claims, and mechanism, formulation, and modality claims. Each covers a distinct layer and can be held by a different owner. The E3-binder and composite-molecule layers are especially decisive.
Has any PROTAC been approved? Yes. In May 2026, vepdegestrant, an oral PROTAC estrogen-receptor degrader developed by Arvinas and Pfizer, received US FDA approval, becoming the first approved PROTAC therapy. Its approval marks the transition of targeted protein degradation from clinical development toward the market. Many other degraders remain in trials.
Why are novel E3 ligases important? Novel E3 ligases are important because the genome encodes more than six hundred E3 ligases but only a few have been harnessed for degradation, so binders for new ligases open a high-value, comparatively uncrowded layer. They can enable tissue- or context-selective degradation and help design around crowded cereblon and von Hippel-Lindau chemistries. Much of the field's future white space lies here.
Where is the white space in targeted protein degradation? The white space includes novel E3-ligase binders, molecular-glue scaffolds and rational glue design, tissue- and ligase-selective degraders, orally bioavailable degraders, and expansion beyond oncology into immunology and neuroscience. The cereblon and von Hippel-Lindau chemistries are comparatively crowded. The durable, defensible value is in new E3 chemistries and glues.
Why does TPD analysis need scientific literature? TPD analysis needs scientific literature because new E3 binders, glue scaffolds, and degradation mechanisms appear in research before they are patented, so the literature gives the earliest signal. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
What software helps analyze the targeted protein degradation patent landscape? Software for the TPD landscape should resolve academic and commercial owners and license chains to canonical entities, cluster the E3-binder, target-ligand, linker, and composite-molecule layers, search patents and scientific literature semantically, and monitor deals and new filings continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Endnotes
- Cowan, A. D., & Ciulli, A. (2022). Driving E3 ligase substrate specificity for targeted protein degradation: lessons from nature and the laboratory. Annual Review of Biochemistry, 91. https://doi.org/10.1146/annurev-biochem-032620-104421
- Fasching, B., Gaínza, P., Oleinikovas, V., Thomä, N. H., et al. (2023). From thalidomide to rational molecular glue design for targeted protein degradation. Annual Review of Pharmacology and Toxicology, 63. https://doi.org/10.1146/annurev-pharmtox-022123-104147
- Ishida, T., & Ciulli, A. (2020). E3 ligase ligands for PROTACs: how they were found and how to discover new ones. SLAS Discovery, 26(4). https://doi.org/10.1177/2472555220965528
- Poongavanam, V., et al. (2024). Molecular glue chemical space and design. Drug Discovery Today. https://doi.org/10.1016/j.drudis.2024.104205
- Zhang, X., et al. (2025). The expanding E3 ligase-ligand landscape for PROTAC technology. Targets, 3(4). https://doi.org/10.3390/targets3040030
- Belcher, B. P., Ward, C. C., & Nomura, D. K. (2021). Ligandability of E3 ligases for targeted protein degradation applications. Biochemistry, 62(3). https://doi.org/10.1021/acs.biochem.1c00464
- Urbina, F., Robertson, N., Hallatt, A. J., & Ciulli, A. (2025). A patent review of von Hippel-Lindau (VHL)-recruiting chemical matter (2019–present). Expert Opinion on Therapeutic Patents, 35(3). https://doi.org/10.1080/13543776.2024.2446232
- Arvinas, Inc. (2026, May 1). Arvinas announces FDA approval of VEPPANU (vepdegestrant) for the treatment of ESR1m, ER+/HER2- advanced breast cancer. https://ir.arvinas.com/news-releases/news-release-details/arvinas-announces-fda-approval-veppanu-vepdegestrant-treatment
- Arvinas, Inc. (2025, September 17). Arvinas provides update on collaboration with Pfizer and announces further actions to support value creation. https://ir.arvinas.com/news-releases/news-release-details/arvinas-provides-update-collaboration-pfizer-and-announces
- Monte Rosa Therapeutics, Inc. (2025, September 15). Monte Rosa Therapeutics announces collaboration with Novartis for degraders to treat immune-mediated diseases. https://ir.monterosatx.com/news-releases/news-release-details/monte-rosa-therapeutics-announces-collaboration-novartis

Semantic search for patents retrieves documents by modeled meaning rather than by lexical overlap. A Boolean or keyword query matches on the surface form of a string: it returns patents whose text contains the specified tokens, combined through operators and often expanded with truncation, proximity, and classification filters. A semantic query matches on representation: each document and the query are encoded as high-dimensional vectors, and retrieval ranks documents by vector similarity, so patents whose claims and disclosures are conceptually related are returned even when they share no keywords with the query. This distinction is decisive in patent work because equivalent subject matter is routinely described in divergent vocabulary, and applicants draft claims with deliberately broad, idiosyncratic, or coined terminology to widen scope. When relevant patents use language the searcher did not anticipate, lexical retrieval fails to surface them, and those recall failures are the primary source of risk in prior art and freedom-to-operate analysis.
The stakes are set by scale. The World Intellectual Property Organization reported roughly 3.5 million patent applications filed worldwide in 2023,¹ and scientific literature has grown at approximately 8 to 9 percent per year across recent decades,² so the candidate space a searcher must cover exceeds what manually constructed Boolean queries can reliably span. Lexical retrieval trades recall for precision: it returns documents matching the exact terms and omits everything phrased differently. Evaluations in the patent-retrieval literature have found that keyword and Boolean logic disregard the syntax and semantics of technical language, which makes patent retrieval materially harder than general-domain information retrieval and leaves relevant documents unretrieved.³,⁴ In prior art and freedom-to-operate, the cost of a single missed document is high, because one overlooked reference can defeat a novelty position or expose a product to infringement liability.
Semantic search also underpins the shift toward agentic AI in R&D and IP. AI agents increasingly query patent and scientific corpora through an API and execute multi-step retrieval and analysis, and dense semantic retrieval is what allows an agent to locate relevant documents without a human hand-crafting Boolean strings. The Model Context Protocol (MCP), introduced by Anthropic in November 2024 and donated to the Linux Foundation's Agentic AI Foundation in December 2025, is now the common standard through which agents connect to external data.⁵ Semantic retrieval is the layer that grounds these agentic workflows in real documents.
How semantic search works
Semantic search rests on learned vector representations. A transformer-based language model, typically fine-tuned on patent and scientific text, encodes each document into a dense embedding, a fixed-length vector of several hundred dimensions whose geometry captures semantic relationships. The query is encoded by the same model, and relevance is scored by a similarity function, most commonly cosine similarity or inner product, over the shared vector space. Because encoding is done offline, retrieval at query time reduces to a nearest-neighbor search: the system returns the documents whose embeddings lie closest to the query embedding, which are the documents most similar in modeled meaning rather than in wording.
At corpus scale, exhaustive comparison against every vector is infeasible, so semantic search relies on approximate nearest-neighbor (ANN) indexing, using structures such as hierarchical navigable small-world graphs or inverted-file quantization to return near-optimal neighbors in sub-linear time. The retrieval model itself is usually a bi-encoder, which embeds query and document independently for speed. A second, more expensive cross-encoder is often applied as a reranking stage over the top candidates, jointly attending to query and document to refine the ordering. This retrieve-then-rerank architecture is standard because it combines the throughput of ANN retrieval with the precision of pairwise scoring.
Dense retrieval demonstrably outperforms lexical baselines. Dense passage retrieval improved top-20 retrieval accuracy by 9 to 19 percentage points over a strong keyword baseline in open-domain benchmarks,⁶ and patent-specific embedding models have been engineered for this domain: the European Patent Office's SEARCHFORMER uses siamese transformer encoders to produce semantic patent embeddings purpose-built for prior art search,⁷ and subsequent work has shown that few-shot fine-tuning and quantized embeddings can adapt these models efficiently to patent retrieval benchmarks.⁸ Three domain-specific factors make this adaptation necessary. Patents are long and structurally heterogeneous, so documents are segmented into passages, and claims are frequently indexed separately because claim language, not the abstract, defines legal scope. Patent corpora are multilingual, so cross-lingual embeddings allow a query in one language to retrieve prior art in another. And patents carry structured metadata, so production systems typically use hybrid retrieval, fusing dense semantic scores with lexical signals and Cooperative Patent Classification or International Patent Classification codes through rank-fusion methods to combine conceptual recall with exact-match precision.
Semantic search becomes substantially more powerful when the vector layer is combined with an explicit knowledge structure. An ontology that formalizes how technologies, materials, methods, and claims relate lets a platform interpret a query within a technology domain and cluster results by concept rather than surface wording. This combination, dense retrieval organized by an ontology, is what separates a modern patent-analytics platform from a keyword database with a search box: it improves recall by retrieving conceptually related work, and it improves interpretability by organizing results into a navigable conceptual structure. Retrieval quality in this setting is measured with recall-oriented metrics such as recall@k and mean average precision, which reflect the priority of finding all relevant documents rather than only the top few.
Where semantic search changes patent work
Prior art search. Prior art search is recall-bound, because the objective is to surface any earlier disclosure bearing on novelty or obviousness. Dense semantic retrieval surfaces prior art expressed in different terminology from the invention, including non-patent literature, which lexical search omits, producing a more complete novelty assessment.
Freedom-to-operate. Freedom-to-operate depends on identifying active, in-force claims that a product could read on, including claims drafted to cover a concept broadly. Semantic retrieval locates relevant claims independent of exact terminology and, combined with claim-level indexing, narrows the analysis to the independent claims that define infringement scope, reducing the coverage gaps that generate FTO risk.
White space analysis. White space analysis depends on clustering activity by concept to expose genuine gaps. Embedding-based clustering distinguishes real conceptual sparsity from apparent sparsity that is only an artifact of divergent terminology, so the identified white space reflects unclaimed technical territory rather than a vocabulary mismatch.
Technology and competitive intelligence. Characterizing a technology area requires connecting related work across patents and scientific literature. Encoding both sources in a shared vector space lets a platform retrieve and align conceptually related documents across them, supporting attribution of activity to technology areas and organizations.
Where Cypris fits
Cypris applies semantic search across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology is what makes the dense retrieval interpretable: it maps how technologies, claims, and research relate, so retrieval is organized by concept rather than surface wording, and results are returned as a navigable conceptual structure rather than a flat ranked list. This lets Cypris surface conceptually relevant patents and papers for prior art, freedom-to-operate at the claim level, and white space analysis, closing the recall gaps that lexical search leaves. Cypris Q, the platform's agentic layer, lets teams run semantic queries conversationally and chain them into multi-step retrieval and analysis, and Agentic Monitoring keeps results current by tracking a technology area over time. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, so AI agents can execute semantic retrieval across the corpus programmatically, and it is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is semantic search for patents?
Semantic search for patents retrieves patents and scientific papers by modeled meaning rather than by exact keywords. It encodes documents and queries as dense vectors and ranks results by vector similarity, so conceptually related documents are returned even when they share no keywords. This closes the recall gaps that cause missed prior art and freedom-to-operate risk in Boolean search.
How is semantic search different from keyword search?
Semantic search differs from keyword search in what it matches. Keyword and Boolean search match the surface form of a query string, while semantic search matches learned vector representations of meaning. For patents this matters because equivalent subject matter is described in divergent vocabulary, and lexical search misses documents phrased differently or drafted with deliberately broad claim language.
What are embeddings in semantic patent search?
Embeddings in semantic patent search are dense vectors, produced by a transformer language model, that encode the meaning of a patent, a claim, a passage, or a query into a shared high-dimensional space. Similarity between embeddings, typically cosine similarity, measures conceptual relatedness. Retrieval returns the documents whose embeddings are nearest to the query embedding.
What is dense retrieval and how does it compare to BM25?
Dense retrieval encodes queries and documents as learned vectors and ranks by vector similarity, whereas BM25 is a sparse, term-frequency lexical method. Dense passage retrieval has been shown to improve top-20 retrieval accuracy by 9 to 19 percentage points over a strong lexical baseline. Production patent systems often combine the two in hybrid retrieval to gain both conceptual recall and exact-match precision.
Why does semantic search matter for prior art search?
Semantic search matters for prior art search because prior art is recall-bound, and relevant disclosures are frequently phrased differently from the invention or appear in non-patent literature. Lexical search omits these, leaving gaps in the novelty assessment. Dense semantic retrieval surfaces conceptually related disclosures regardless of wording, making the assessment more complete.
How does semantic search improve freedom-to-operate analysis?
Semantic search improves freedom-to-operate analysis by locating active claims a product could read on even when those claims use different terminology or cover a concept broadly. Combined with claim-level indexing, it focuses the analysis on the independent claims that define infringement scope. This reduces the coverage gaps that are the main source of FTO risk.
What is a retrieve-then-rerank pipeline?
A retrieve-then-rerank pipeline is a two-stage architecture. A fast bi-encoder retrieves a candidate set using approximate nearest-neighbor search, then a more expensive cross-encoder rescores the top candidates by jointly attending to the query and each document. This combines the throughput of vector retrieval with the precision of pairwise relevance scoring.
Does semantic search need an ontology?
Semantic search does not strictly require an ontology, but combining the two is substantially more powerful. An ontology formalizes how technologies relate, letting a platform interpret a query within a domain and cluster results by concept. Cypris combines dense semantic retrieval with a proprietary R&D ontology across more than 500 million patents and scientific papers for this reason.
How do AI agents use semantic search for patents?
AI agents use semantic search as the retrieval layer that lets them locate relevant patents and papers without a human writing Boolean strings, enabling autonomous multi-step analysis. Agents query the corpus through an API and use dense retrieval to ground their reasoning in real documents. Cypris offers enterprise API partnerships with OpenAI, Anthropic, and Google so agents can run semantic retrieval across its corpus.
Which teams benefit from semantic patent search?
Semantic patent search benefits R&D, innovation, and IP teams running prior art, freedom-to-operate, white space, and technology-intelligence searches, all of which depend on high-recall retrieval of conceptually relevant documents. It is most valuable in research-intensive industries such as pharmaceuticals, chemicals, advanced materials, and energy. Cypris serves hundreds of enterprise customers across these industries.
Endnotes
- World Intellectual Property Organization. World Intellectual Property Indicators (annual series). https://www.wipo.int/publications/
- Bornmann, L. & Mutz, R. (2015). Growth rates of modern science: a bibliometric analysis based on the number of publications and cited references. Journal of the Association for Information Science and Technology. https://doi.org/10.1002/asi.23329
- Zihayat, M. & Etwaroo, R. (2021). A non-factoid question answering system for prior art search. Expert Systems with Applications. https://doi.org/10.1016/j.eswa.2021.114910
- Lupu, M., Piroi, F., Hanbury, A. & Zenz, V. (2011). CLEF-IP 2011: Retrieval in the Intellectual Property Domain.
- Anthropic (2025). Donating the Model Context Protocol and establishing the Agentic AI Foundation. https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation; Linux Foundation (2025). Formation of the Agentic AI Foundation. https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation
- Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP. https://doi.org/10.18653/v1/2020.emnlp-main.550
- Vowinckel, K. & Hähnke, V. D. (2023). SEARCHFORMER: Semantic patent embeddings by siamese transformers for prior art search. World Patent Information. https://doi.org/10.1016/j.wpi.2023.102192
- Chikkamath, R. et al. (2025). Patent Retrieval with Few-Shot Fine-Tuning and Quantized Embeddings. https://doi.org/10.1145/3787279.3787295
