A powerful new foundation for custom queries—built on Lucene and designed for R&D precision.
Over the past few years, Cypris has helped innovation teams make faster, more informed decisions by centralizing critical insights across datasets like patents, academic papers, and company activity. But until now, our search experience relied on a legacy query system with limited capabilities, offering little support for advanced search features or dataset-level customization.
Today, we’re excited to introduce an upgraded Advanced Search on Cypris, a complete overhaul of our query engine and search experience, powered by the open-standard Lucene query syntax. This update introduces a more robust and flexible search foundation, unlocking new ways to query data, build complex filters, and extract precisely what you need across patents, research, and more.
Why we rebuilt our search system from the ground up
Cypris’ original query syntax, a proprietary format used internally for years, limited users’ ability to craft advanced queries or tailor searches to specific datasets. It lacked modern capabilities like proximity searches, field-level customization, or true Boolean logic. This made it difficult to build a reliable and intuitive experience for both casual users and advanced researchers.
By moving to Lucene, we’re adopting a powerful, industry-standard query language that makes it easier for developers to build advanced features—and gives users access to a far more capable and flexible search toolset.
What’s new in Advanced Search
1. Custom Queries by Dataset
You can now layer queries to search across datasets or tailor filters to each one. For example, you can run a broad query on drone delivery, and then add separate layers to focus on patents by a specific assignee and papers from a specific country or funding agency.
Navigating the All Datasets tab introduces a new level of complexity—and power—by allowing users to apply dataset-specific logic within a single, unified query workflow. While querying multiple datasets simultaneously might seem straightforward, the underlying differences in schema, metadata, and available fields between our proprietary datasets make this a deeply technical challenge. Patents, for example, include claims, application numbers, and multiple date fields (filed, granted, updated), while academic papers use DOIs, have different structural conventions, and emphasize different metadata. In the past, we sidestepped this complexity by translating general queries like ((drone_allText)) into dataset-specific logic under the hood. Now, instead of obscuring that logic, we allow users to opt in to it. The builder provides progressive layers of customization: start with intuitive keyword searches across all fields, then move into the advanced builder for field-specific targeting, fuzzy logic, and term boosting, and finally, tailor query logic by dataset—such as specifying different countries of interest for papers vs. patents. This approach preserves flexibility while giving users full control, and with tools like our real-time Live Analysis and “Your Query” panel, we make it easy to understand how every decision affects the results.
2. More Fields to Query
We’re exposing deeper fields across datasets—giving you explicit control over the dimensions of your search. For the first time, users can now search academic papers by DOI, a critical identifier previously unsupported on the platform. You can also query by:
- Author or inventor names
- Organizations or assignees
- Countries, journals, funding agencies, and more
3. Full Boolean Support
Advanced Search now leverages powerful Boolean logic—AND, OR, NOT, and grouping—enabling more precise control over search logic and improving performance and accuracy.
4. Lucene Syntax Features
Use built-in Lucene features to create expressive, complex searches:
- Proximity searches to find terms near each other
- Fuzzy searches for flexible matching
- Exact phrase matching
- Boosting to prioritize results (e.g., prioritize results mentioning AI 3x more than others)
- Prefix/Postfix queries to match phrases that start or end a certain way
- Range queries for fields like date, funding amounts, or numerical values
A more powerful user experience
Our new search interface is built to help you tap into these capabilities without needing to know the syntax from the start. You’ll find:
- A Query Builder to guide you through complex searches
- A Help Video to onboard users to Lucene-style searches
- Inline examples and tips for writing queries using grouping, boosting, and more
Built for precision, speed, and customization
With Lucene as our foundation, search results are now not only more flexible but also faster and more accurate. Semantic search continues to offer natural-language ease of use, while Boolean search gives power users the performance and structure they need to uncover insights with greater specificity.
Whether you’re an innovation analyst drilling into AI patents or a business development lead scanning academic papers from Chilean researchers—Advanced Search is built to help you get to the signal, faster.
Available now to all users
Advanced Search is live and available across the Cypris platform today. If you’re already using Cypris, you’ll find the new search interface in your dashboard, complete with updated syntax documentation and walkthroughs.
We’re excited to see what you’ll build, discover, and analyze with this new capability. This is just the beginning—we’ll continue expanding the fields, syntax features, and customization options as we push the boundaries of what intelligent search can do for R&D.

Introducing Advanced Search on Cypris

A powerful new foundation for custom queries—built on Lucene and designed for R&D precision.
Over the past few years, Cypris has helped innovation teams make faster, more informed decisions by centralizing critical insights across datasets like patents, academic papers, and company activity. But until now, our search experience relied on a legacy query system with limited capabilities, offering little support for advanced search features or dataset-level customization.
Today, we’re excited to introduce an upgraded Advanced Search on Cypris, a complete overhaul of our query engine and search experience, powered by the open-standard Lucene query syntax. This update introduces a more robust and flexible search foundation, unlocking new ways to query data, build complex filters, and extract precisely what you need across patents, research, and more.
Why we rebuilt our search system from the ground up
Cypris’ original query syntax, a proprietary format used internally for years, limited users’ ability to craft advanced queries or tailor searches to specific datasets. It lacked modern capabilities like proximity searches, field-level customization, or true Boolean logic. This made it difficult to build a reliable and intuitive experience for both casual users and advanced researchers.
By moving to Lucene, we’re adopting a powerful, industry-standard query language that makes it easier for developers to build advanced features—and gives users access to a far more capable and flexible search toolset.
What’s new in Advanced Search
1. Custom Queries by Dataset
You can now layer queries to search across datasets or tailor filters to each one. For example, you can run a broad query on drone delivery, and then add separate layers to focus on patents by a specific assignee and papers from a specific country or funding agency.
Navigating the All Datasets tab introduces a new level of complexity—and power—by allowing users to apply dataset-specific logic within a single, unified query workflow. While querying multiple datasets simultaneously might seem straightforward, the underlying differences in schema, metadata, and available fields between our proprietary datasets make this a deeply technical challenge. Patents, for example, include claims, application numbers, and multiple date fields (filed, granted, updated), while academic papers use DOIs, have different structural conventions, and emphasize different metadata. In the past, we sidestepped this complexity by translating general queries like ((drone_allText)) into dataset-specific logic under the hood. Now, instead of obscuring that logic, we allow users to opt in to it. The builder provides progressive layers of customization: start with intuitive keyword searches across all fields, then move into the advanced builder for field-specific targeting, fuzzy logic, and term boosting, and finally, tailor query logic by dataset—such as specifying different countries of interest for papers vs. patents. This approach preserves flexibility while giving users full control, and with tools like our real-time Live Analysis and “Your Query” panel, we make it easy to understand how every decision affects the results.
2. More Fields to Query
We’re exposing deeper fields across datasets—giving you explicit control over the dimensions of your search. For the first time, users can now search academic papers by DOI, a critical identifier previously unsupported on the platform. You can also query by:
- Author or inventor names
- Organizations or assignees
- Countries, journals, funding agencies, and more
3. Full Boolean Support
Advanced Search now leverages powerful Boolean logic—AND, OR, NOT, and grouping—enabling more precise control over search logic and improving performance and accuracy.
4. Lucene Syntax Features
Use built-in Lucene features to create expressive, complex searches:
- Proximity searches to find terms near each other
- Fuzzy searches for flexible matching
- Exact phrase matching
- Boosting to prioritize results (e.g., prioritize results mentioning AI 3x more than others)
- Prefix/Postfix queries to match phrases that start or end a certain way
- Range queries for fields like date, funding amounts, or numerical values
A more powerful user experience
Our new search interface is built to help you tap into these capabilities without needing to know the syntax from the start. You’ll find:
- A Query Builder to guide you through complex searches
- A Help Video to onboard users to Lucene-style searches
- Inline examples and tips for writing queries using grouping, boosting, and more
Built for precision, speed, and customization
With Lucene as our foundation, search results are now not only more flexible but also faster and more accurate. Semantic search continues to offer natural-language ease of use, while Boolean search gives power users the performance and structure they need to uncover insights with greater specificity.
Whether you’re an innovation analyst drilling into AI patents or a business development lead scanning academic papers from Chilean researchers—Advanced Search is built to help you get to the signal, faster.
Available now to all users
Advanced Search is live and available across the Cypris platform today. If you’re already using Cypris, you’ll find the new search interface in your dashboard, complete with updated syntax documentation and walkthroughs.
We’re excited to see what you’ll build, discover, and analyze with this new capability. This is just the beginning—we’ll continue expanding the fields, syntax features, and customization options as we push the boundaries of what intelligent search can do for R&D.

Keep Reading

Chemical and enzymatic recycling has become the frontier of the circular economy for plastics, and its patent landscape is distinctive because molecular recycling is not one technology but a set of competing routes, each with its own chemistry, feedstocks, and process engineering. Where mechanical recycling melts and reforms plastic, losing quality with each cycle and struggling with colored, mixed, or contaminated waste, molecular recycling breaks polymers back down to their building blocks, monomers or feedstock chemicals, that can be repurified and repolymerized to a quality equivalent to virgin material. The routes divide by polymer and mechanism: engineered enzymes depolymerize polyester and PET even when colored or contaminated, with machine-learning-aided enzyme engineering producing fast, robust variants;¹ chemical routes such as methanolysis, glycolysis, and hydrolysis cleave PET into its monomers; and pyrolysis and gasification convert polyolefins such as polyethylene and polypropylene into oils and feedstocks. A persistent constraint on the enzymatic route is that highly crystalline PET resists depolymerization, so pretreatment and reaction-medium engineering matter as much as the enzyme itself,³ and thermostable, efficient enzymes remain difficult to engineer.⁴,⁵ Because each route is a distinct region of patenting, freedom-to-operate and white space analysis must treat plastic recycling as several landscapes at once, spanning enzymes and catalysts, reactor and process design, feedstock pretreatment, and monomer purification.
The field is being pulled forward by regulation and by the arrival of commercial-scale plants. In the European Union, the Packaging and Packaging Waste Regulation sets binding minimum-recycled-content targets for plastic packaging for 2030 and 2040 and requires packaging to be recyclable by 2030, creating durable demand for high-quality recycled material that mechanical recycling cannot fully supply.⁶ At the same time, enzyme discovery is industrializing: a systematic profiling of PET-depolymerizing enzymes mapped roughly 1,894 candidate hydrolases across about 170 natural sequence lineages, vastly expanding the toolbox beyond the handful of enzymes studied a few years ago.² Commercial facilities are moving from demonstration to build-out, with a first industrial enzymatic PET plant designed for about 50,000 tonnes per year of prepared post-consumer waste, on a revised commissioning schedule.⁷ The patent record reflects both the biology and the chemistry: across the Cypris corpus of more than 500 million patents and scientific papers, the PET-depolymerization and molecular-recycling set holds on the order of 1,266 families and grew from about 26 in 2020 to roughly 171 in 2024, with the most active assignees mixing chemical-route incumbents such as IFP Energies Nouvelles and Eastman with enzymatic and brand-side players such as Jeplan and Coca-Cola, and the United States, France, and China leading on geography; 2025 and 2026 counts are partial because of the publication lag.
The strategic question is which route and polymer to back, and the white space sits where the chemistry is hardest. Engineered enzymes have advanced furthest for PET and polyester, so the open, high-value ground is increasingly in enzymes and processes for other polymers, in the machine-learning-guided discovery and engineering of new depolymerizing enzymes,²,⁵ and in handling mixed and contaminated feedstocks. Chemical routes for PET are maturing, while the chemical recycling of polyolefins, the largest share of plastic waste, remains harder and less crowded, particularly in upgrading pyrolysis oils to usable feedstocks. Textile-to-textile recycling of polyester is a further emerging layer. Reading the landscape by route, polymer, and process, and tracking both the patents and the underlying enzyme and catalysis research, is what separates a crowded region from an open one.
Where the plastic-recycling white space is
Enzymes for non-PET polymers. Engineered enzymes that depolymerize polyolefins, nylons, and polyurethanes, rather than only polyester, are an early, high-value target as enzymatic PET matures.
Machine-learning enzyme discovery. Computational discovery and engineering of new, more thermostable and efficient depolymerases is a fast-moving layer that compresses development time.²,⁵
Polyolefin chemical recycling. Pyrolysis and its upgrading to usable feedstocks for the largest category of plastic waste remain harder and less crowded than PET routes.
Mixed and contaminated feedstocks. Processes that handle colored, multilayer, and mixed waste, which mechanical recycling cannot, are a differentiating capability.
Monomer purification and textile recycling. High-purity monomer recovery and textile-to-textile polyester recycling are distinct, emerging layers where quality and economics are decided.
How AI-powered landscape and white space analysis helps
Resolving a landscape that spans enzymatic and chemical routes across several polymers, under a regulatory recycled-content pull, requires more than keyword search. AI-powered analysis addresses this with semantic search that clusters activity by route, polymer, and process across varied terminology, attribution that normalizes filers to canonical entities and tracks new entrants, and continuous monitoring that keeps pace with a fast-commercializing field. Because recycling advances appear in scientific and enzyme-engineering literature before they are patented, reading both patents and literature gives the earliest signal of where scalable routes are emerging.
Where Cypris fits
Cypris runs patent landscape and white space analysis for multi-route fields such as chemical and enzymatic plastic recycling across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by route, enzymatic, methanolysis, glycolysis, hydrolysis, and pyrolysis, and by polymer and process layer, and normalizes filers to canonical entities, so a team can resolve which routes and polymers are crowded and which remain open as white space, and can track new entrants as the field scales. Semantic search across patents and scientific literature connects filings to the underlying enzyme-engineering and catalysis research, which is where recycling advances appear first. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined route over time and flags new patents and papers as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is chemical and enzymatic plastic recycling? Chemical and enzymatic plastic recycling, also called molecular recycling, breaks plastics back down to monomers or feedstock chemicals that can be repolymerized to virgin quality, unlike mechanical recycling, which degrades quality. Routes include engineered-enzyme depolymerization, methanolysis, glycolysis, and hydrolysis for PET, and pyrolysis and gasification for polyolefins. Each is a distinct region of patenting.
Why is molecular recycling a patenting hotspot? Molecular recycling is a patenting hotspot because recycled-content rules and packaging regulations are creating durable demand for high-quality recycled material, and the first commercial plants are coming online. It can process colored, mixed, and contaminated waste that mechanical recycling cannot. That combination is driving filings across enzymes, catalysts, and processes.
What routes does the plastic-recycling landscape cover? The landscape covers engineered-enzyme depolymerization of polyester and PET, chemical depolymerization of PET by methanolysis, glycolysis, and hydrolysis, and pyrolysis and gasification of polyolefins. Each route has distinct enzymes or catalysts, reactors, and purification steps. Freedom-to-operate and white space analysis must treat them separately.
Where is the white space in plastic recycling? The white space includes engineered enzymes for non-PET polymers, machine-learning-guided enzyme discovery, polyolefin chemical recycling and pyrolysis-oil upgrading, processes for mixed and contaminated feedstocks, and monomer purification and textile-to-textile recycling. Enzymatic PET is comparatively advanced. The most open, high-value opportunities are in other polymers and in polyolefin routes.
Why is polyolefin recycling harder than PET recycling? Polyolefin recycling is harder because polyethylene and polypropylene lack the cleavable bonds that make PET amenable to enzymatic and chemical depolymerization, so they are typically broken down by pyrolysis into mixed oils that must then be upgraded. Polyolefins are also the largest share of plastic waste. That difficulty leaves the layer less crowded and high in value.
Why does plastic-recycling analysis need scientific literature? Plastic-recycling analysis needs scientific literature because enzyme-engineering and catalysis advances appear in research before they are patented, so the literature gives the earliest signal. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
What software helps analyze the plastic-recycling patent landscape? Software for the plastic-recycling landscape should cluster activity by route, polymer, and process, resolve filers to canonical owners, search patents and scientific literature semantically, and monitor a fast-commercializing field continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams use plastic-recycling patent landscape analysis? Plastic-recycling patent landscape analysis is used by R&D, innovation, IP, and strategy teams at chemicals, materials, packaging, and waste-management companies, enzyme and catalysis developers, and their partners, as well as investors and policymakers. It informs which route to back, where to file, and where competitors are concentrated. Cypris serves hundreds of enterprise customers across chemicals, advanced materials, energy, and other regulated industries.
Endnotes
- Lu, H., Diaz, D. J., Shroff, P., Alper, H. S., et al. (2022). Machine learning-aided engineering of hydrolases for PET depolymerization. Nature, 604. https://doi.org/10.1038/s41586-022-04599-z
- Park, S. Y., Ki, D., Sagong, H.-Y., Seo, H., et al. (2025). Landscape profiling of PET depolymerases using a natural sequence cluster framework. Science, 387(6729). https://doi.org/10.1126/science.adp5637
- Kaabel, S., Therien, J. P. D., Auclair, K., et al. (2021). Enzymatic depolymerization of highly crystalline PET enabled in moist-solid reaction mixtures. Proceedings of the National Academy of Sciences, 118(29). https://doi.org/10.1073/pnas.2026452118
- Kawabata, T., Iizuka, R., & Kawai, F. (2024). Engineered polyethylene terephthalate hydrolases: perspectives and limits. Applied Microbiology and Biotechnology, 108. https://doi.org/10.1007/s00253-024-13222-2
- Li, Q., Su, L., Dian, L., et al. (2024). Dynamic docking-assisted engineering of hydrolases for efficient PET depolymerization. ACS Catalysis, 14(9). https://doi.org/10.1021/acscatal.4c00400
- European Commission, Directorate-General for Environment. Packaging and packaging waste. https://environment.ec.europa.eu/topics/waste-and-recycling/packaging-waste_en
- Carbios. PET biorecycling technology. https://www.carbios.com/en/pet-biorecycling-technology/

The fastest way to turn a commodity AI assistant into a reliable R&D and IP research tool is to connect it to a domain-oriented intelligence layer through the Model Context Protocol, because the general-purpose model supplies the reasoning while the verticalized agent supplies the grounded, high-signal data the model cannot hold on its own. This is the single architectural decision that separates an AI that drafts plausible-sounding patent summaries from one an innovation team can actually act on. The model you start with is a commodity. The vertical integration you attach to it is the differentiator.
This guide explains what commodity AI gets wrong in R&D and IP work, why the gap is structural rather than a matter of prompting, and how a domain MCP integration closes it. It is written for R&D directors, IP managers, and innovation strategists who already have access to capable general models and want to understand what it takes to make them trustworthy for stage-gate decisions.
What Commodity AI Means in an R&D Context
A commodity AI is a general-purpose large language model accessed through a chat interface or an enterprise assistant, the same model available to every competitor in your market. These horizontal systems are built on broad pre-training across diverse public data and are designed to handle a wide range of tasks without deep subject knowledge [1]. They are genuinely useful for summarizing a document you paste in, drafting an email, or explaining a concept. The strength of the horizontal model is breadth and speed of deployment.
The weakness is that breadth is the wrong shape for R&D and IP intelligence. A prior art search, a freedom-to-operate question, or a white space analysis does not reward general fluency. It rewards completeness, recency, and precision against a defined corpus of patents and scientific literature. A commodity model has no live connection to that corpus. It answers from a frozen snapshot of training data and from whatever you happened to paste into the prompt, which means the most consequential R&D questions are exactly the ones it is least equipped to answer.
Why the Gap Is Structural, Not a Prompting Problem
The instinct when a general model gives a weak patent answer is to write a better prompt. This helps at the margin, but it cannot solve the core problem, because the failure is rooted in two structural limits that prompting does not touch.
The first limit is hallucination. Generating plausible but ungrounded output remains the single biggest barrier to deploying language models in production as of 2026, and complete elimination is not possible because the tendency is tied to the model's generative capability itself [2]. In an IP context this is not a cosmetic flaw. A model conducting an ungrounded prior art search can surface references that do not exist, misattribute a claim, or describe a system that is physically impossible, and it delivers all of it in the same confident register as a correct answer [3]. A 2026 study evaluating five popular public models on preliminary prior art searches found that accuracy, consistency, and the ability to surface conceptually relevant art from adjacent fields varied widely and required careful human verification [4]. The authority of the output is not evidence of its reliability.
The second limit is that flooding a general model with more data does not fix the first problem and often makes it worse. There is a temptation to solve grounding by dumping an entire patent dataset into the model's context window. Research on context engineering shows this backfires. As a broad, undifferentiated corpus fills the context window, the model's ability to reason over it degrades, an effect documented across multiple studies of how models use long contexts [5][6]. The model does not get smarter as you add data. Past a point, it gets less accurate. This is why raw access to a large dataset is not the same as intelligence over it, and why the path to reliability runs through retrieving the right small set of high-signal documents rather than the largest possible set.
Together these two limits define the gap. The commodity model is fluent but ungrounded, and you cannot ground it simply by giving it everything. You ground it by connecting it to a system that already knows which fraction of the corpus matters for the question being asked.
What a Verticalized Agent Adds
A vertical AI agent is purpose-built for a specific domain, pre-loaded with domain knowledge, proprietary data models, and deep integrations into the systems where that domain's data lives [7]. Where a horizontal agent relies on broad pre-training, a vertical agent demands domain adaptation and plugs into domain-specific data pipelines, and it is this depth that produces superior accuracy, compliance, and reliability within its field [1]. The market has moved decisively in this direction. Industry analysts forecast that vertical-first deployments will account for a large and growing share of enterprise AI in 2026, with industry-specific AI solutions growing far faster than general-purpose tools, because the highest-return deployments come from embedding agents into existing domain workflows rather than buying a generic assistant [8].
In R&D and IP, the domain adaptation that matters is an ontology. A proprietary R&D ontology lets a vertical agent understand that a query about a polymer coating, a thermal barrier, and a specific chemical family are related concepts in a way a keyword search never will, and it lets the agent retrieve the conceptually relevant subset of patents and papers rather than a lexical match. That is the precise capability the commodity model lacks and the precise reason it cannot be prompted into existence. The ontology is the difference between access to 500 million patents and scientific papers and intelligence over them.
Where MCP Fits
The Model Context Protocol is the open standard that lets a general model call an external system as a tool during a conversation, which is what makes the upgrade from commodity AI to verticalized agent a connection rather than a rebuild [9]. You do not have to abandon the general model your team already uses. MCP is the mechanism by which that model reaches out, mid-reasoning, to a domain-oriented layer, asks it a scoped question, and receives back a reasoned, grounded answer rather than a raw dump of records.
This is the architectural pattern that resolves the structural gap. The general model continues to do what it is good at, which is language, synthesis, and conversation. The vertical agent does what it is good at, which is retrieving the high-signal subset from a defined corpus and reasoning within the domain. The protocol connects them. Crucially, because the vertical layer returns a scoped and reasoned result rather than the entire dataset, it sidesteps the context degradation problem entirely. The model never has to hold the full corpus in its context window, so its reasoning stays sharp.
How the Upgrade Works in Practice
The practical sequence is straightforward to describe even though the engineering behind the vertical layer is substantial. A researcher asks a question in the AI interface they already use. The general model recognizes that the question requires domain intelligence and, through MCP, routes a scoped query to the domain-oriented R&D layer. That layer uses its ontology to retrieve the relevant patents and scientific papers, reasons over them within the domain, and returns a grounded finding. The general model then composes that finding into a clear answer for the researcher. The researcher experiences one fluid conversation. Underneath it, the work has been divided between the part of the system built for language and the part built for the domain.
This division maps directly onto the R&D and IP stage-gate process. A prior art agent built this way returns grounded references rather than invented ones. A white space analysis returns a defensible read of where the unclaimed territory sits. A freedom-to-operate question is answered against live patent data rather than a stale training snapshot. Regulatory tracking stays current because the vertical layer, not the frozen model, is the source of truth. In each case the commodity model is the interface and the verticalized agent is the engine.
What This Means for Buyers
The strategic takeaway is that the model is no longer where the advantage lives. Every competitor in your market can access the same capable general models, which is precisely what makes them a commodity. The durable advantage comes from what you connect those models to. An organization that wires its general AI to a domain-oriented R&D intelligence layer through MCP gets grounded, current, defensible answers to its most important innovation questions. An organization that relies on the commodity model alone gets fluent guesses. The gap between those two outcomes is not the model. It is the vertical integration.
Cypris is built to be that vertical layer. As an enterprise R&D intelligence platform spanning more than 500 million patents and scientific papers, organized by a proprietary R&D ontology and powered by Cypris Q agentic workflows, it is designed to deliver domain-oriented intelligence to the AI systems R&D and innovation teams already use, through enterprise API partnerships with OpenAI, Anthropic, and Google [10]. Rather than asking a general model to be an IP expert it cannot be, Cypris supplies the grounded domain reasoning the model needs, across the workflows that matter most: prior art agents, white space analysis, freedom-to-operate, and regulatory tracking. The commodity model handles the conversation. Cypris handles the intelligence.
Frequently Asked Questions
What does it mean to upgrade commodity AI with a vertical agent?
It means connecting a general-purpose AI model to a domain-specific intelligence system so the model can answer specialized questions accurately. The general model provides language and reasoning, while the vertical agent provides grounded, high-signal data from a defined corpus such as patents and scientific papers. The connection is what turns a fluent generalist into a reliable domain tool.
Why can't I just use a better prompt to get good patent answers from a general AI?
Prompting helps at the margin but cannot solve the core problem, because the failure is structural. A general model has no live connection to patent and scientific data and answers from a frozen training snapshot, so it can hallucinate references that do not exist. Better prompts cannot create data access the model fundamentally lacks.
What is the Model Context Protocol and why does it matter here?
The Model Context Protocol, or MCP, is an open standard that lets a general AI model call an external system as a tool during a conversation. It matters because it allows a commodity model to reach a domain-oriented intelligence layer mid-reasoning and receive a grounded answer. MCP is the mechanism that connects a general model to a vertical agent without replacing the model.
Won't connecting my AI to a huge patent database make it smarter?
Not on its own. Research on context engineering shows that flooding a model's context window with a broad, undifferentiated corpus degrades its reasoning rather than improving it. The value comes from a system that retrieves the small, high-signal subset relevant to your question, not from raw access to the largest possible dataset.
What is the difference between a horizontal AI agent and a vertical AI agent?
A horizontal agent is general-purpose and built for breadth across many tasks and departments, with broad pre-training and fast deployment. A vertical agent is purpose-built for a single domain, pre-loaded with domain knowledge and integrated into domain-specific data pipelines. Vertical agents take longer to build but deliver superior accuracy and reliability within their field.
Why is hallucination such a serious problem for R&D and IP work?
Because in prior art and freedom-to-operate work, a confident wrong answer can misdirect a real innovation or legal decision. Hallucination remains the biggest barrier to production deployment of language models in 2026, and a model can surface non-existent references in the same authoritative tone as correct ones. The authority of the output is not evidence of its accuracy.
What role does an ontology play in a vertical R&D agent?
An ontology lets the agent understand conceptual relationships between technologies, materials, and methods rather than relying on keyword matching. This allows it to retrieve patents and papers that are conceptually relevant even when they use different terminology. The ontology is the core capability that makes a vertical agent precise where a general model is not.
Do I have to replace my existing AI tools to do this?
No. The entire point of an MCP-based integration is that you keep the general AI your team already uses and connect it to a vertical intelligence layer. The general model remains the interface, and the domain agent works behind it. The upgrade is a connection, not a rebuild.
How does this approach map to my R&D workflow?
It maps directly onto stage-gate work. A prior art agent returns grounded references, a white space analysis returns a defensible read of unclaimed territory, a freedom-to-operate query runs against live patent data, and regulatory tracking stays current through the vertical layer. Each workflow is answered by the domain engine rather than the frozen general model.
If everyone can access the same AI models, where is the competitive advantage?
The advantage is no longer the model, which is exactly why it is a commodity. It comes from what you connect the model to. An organization that wires its general AI to a domain-oriented R&D intelligence layer gets grounded, defensible answers, while one relying on the model alone gets fluent guesses.

For most of the past three decades, the corporate IP team occupied a clear position near the end of the innovation process. Research and development explored a concept, leadership committed resources, scientists and engineers built the product, and only then did the work reach IP for protection, prosecution, and portfolio management. IP was a service function, expert and essential, but downstream of the decisions that mattered most. That sequence has quietly inverted. Today R&D comes to IP before resources are committed, asking what already exists in the patent record and treating the answer as a go or no-go signal on whether to pursue an idea at all. A prior art search is no longer just a legal precaution. It has become a strategic input that shapes which programs get funded, which get redirected, and which get killed before a dollar is spent.
This is a meaningful elevation of the IP team's role, and in most organizations it happened by default rather than by design. The mandate expanded because R&D became too expensive and too risky to pursue on instinct. The data and the tooling underneath the IP function, however, did not expand with it. The team is now being asked forward-looking strategic questions and is answering them with the one dataset it has always owned: the patent record. That mismatch between the question being asked and the data available to answer it is the source of a specific, costly, and underappreciated error. It has a name worth retiring from strategic vocabulary: the white space fallacy, the assumption that an empty region of the patent map is an open opportunity.
The stakes are higher than the tooling reflects
The reason this matters is that the decisions riding on these analyses are enormous, and the base rates for innovation are unforgiving. Failure rates across corporate R&D are persistently high. Industry research has long pegged new product failure somewhere between a third and half of all launches, and a substantial share of R&D projects never reach production at all. These failures have many causes, but a recurring and underexamined one is the practice of validating technical opportunity through patent analysis while leaving commercial opportunity unvalidated. A program clears the patent landscape, looks open, and proceeds, only to discover that the space was empty for reasons the patent record never showed. When the IP team's answer is steering investment direction, the cost of an incomplete map is no longer a missed filing. It is a misallocated research budget and a multi-year bet placed in the wrong direction.
White space and opportunity space are not the same thing
The cleanest way to see the error is to picture two overlapping circles. The first is patent white space, the regions of a technology landscape where few or no active patents exist. The second is commercial opportunity, the areas where genuine market demand and commercial momentum are forming. The portfolio every organization actually wants sits in the overlap, where a defensible technical position meets real commercial pull. That overlap is a narrow slice, and most teams cannot see it clearly because they are looking at only one of the two circles.
The reason patent white space gets mistaken for opportunity is structural rather than careless. Patent data is the dataset the IP team owns, the tool it has on hand, and the answer it can produce on demand. So the strategic question silently narrows from where should we invest to where is the patent map empty, and those two questions only sometimes have the same answer. The narrowing is invisible because it happens inside the framing of the analysis, not in its conclusions. Everyone in the room believes they are discussing opportunity. They are actually discussing patent density.
An empty region of the patent map can mean two very different things, and distinguishing between them is the whole game. It can be open for a reason, because there is no market demand, because the underlying science does not work yet, or because the unit economics never close. Easy to patent does not mean possible to monetize, and a clear space on the map can simply be a place no one has bothered to claim because there is nothing there worth claiming. Alternatively, the empty space can be a trap of the opposite kind, a region where competitors are very much active but moving through channels that never touch the patent system: trade secrets, defensive publications, or simply faster commercial execution that outruns the filing timeline. In both cases the patent map looks identical. It looks open. Only data drawn from outside the patent system can tell you which kind of empty you are actually looking at, and the two demand completely different strategic responses.
The inverse error is just as expensive and far less discussed. Some of the most contested, patent-dense regions of a landscape are exactly where the market is moving, and exactly where a given organization may be dangerously under-protected. A crowded patent map instinctively reads as a closed door, a market already won by incumbents. But density is a measure of competitive intensity, not of whether the opportunity is worth pursuing. Some of the most commercially urgent positions a company can take are in crowded spaces where the organization holds a real technical advantage but has under-filed relative to the competition. Reading crowdedness as a stop sign can forfeit exactly the positions most worth fighting for.
A patent is a twenty-year bet placed with rear-view data
Underneath the white space problem sits a deeper structural mismatch, this one about time. A patent is a roughly twenty-year commitment. That makes it one of the most forward-looking instruments a company holds, a claim staked on what will matter for two decades. Yet the patent record itself is one of the most backward-looking datasets available to anyone. Applications publish around eighteen months after they are filed, and the decisions behind them were made well before that. By the time a filing is visible in the public record, it describes a strategic choice that may be two or three years old. Patents are lagging indicators, sometimes by years, as applications crawl through prosecution. A team that validates a long-horizon investment using only existing patents is steering a twenty-year bet with a dataset that describes where the field was, not where it is going.
The question the IP team is increasingly asked to answer is whether a given portfolio or technology area will still matter in five to ten years. Answering that honestly requires three categories of signal that the patent record either omits entirely or reports too late to be useful.
The first is scientific momentum. Peer-reviewed papers, preprints, grant awards, and clinical activity reveal where the underlying technology is heading long before any of it reaches a patent application. Preprints in particular can surface a competitor's technical direction months to years ahead of the corresponding filing, because the science is published when it is done, not when the legal strategy is finalized. A field rich in recent publication but thin on filings is frequently an emerging opportunity, an early window in which an organization can establish a position before the patent landscape fills in and the easy ground is taken. To a patent-only view, that same field registers as white space and risks being dismissed as empty, when it is in fact the most valuable kind of crowded: crowded with science, not yet with claims.
The second is commercial signal. Venture funding, startup formation, mergers and acquisitions, corporate disclosures, and product launches reveal where commercial conviction is forming, frequently well ahead of patent activity. A technology domain showing minimal patent filings but hundreds of millions of dollars in aggregate venture funding is not white space. It is a market building momentum through channels that patent analytics simply cannot see. When an acquirer buys a startup, the strategic implication for every competitor in the space is immediate, but the patent assignment record may take months to update, and the commercial rationale for the deal, which market is being targeted, which product lines will expand, which competing approaches are being consolidated, never enters the patent data at all. That intelligence lives in deal records, regulatory filings, and corporate disclosures, in a layer of the landscape the patent-only team never sees.
The third is forward indicators, the signals that point at intent before it materializes as anything protectable. Regulatory filings, clinical pipelines, market intelligence, and hiring patterns all belong here. Hiring is among the most underused signals of all. The engineering and research roles a company is staffing frequently describe, in the job specifications themselves, exactly what the organization is building, and they appear long before any of that work surfaces as a filing. A competitor assembling a team around a specific technical capability is making a far earlier and often far clearer statement of direction than anything that will eventually reach a patent office.
None of this argues for abandoning patent data. Global patents remain the foundation, the authoritative record of what has actually been claimed and protected, and no serious analysis proceeds without them. The argument is narrower and harder to dismiss: patents are necessary but not sufficient for the strategic questions IP teams are now expected to answer. The foundation is solid. The problem is that three of the four walls are missing, and the team is being asked to assess the whole structure from the foundation alone.
Why the gap persists when it is so clearly understood
If the gap is this obvious, the fair question is why it endures across so many sophisticated organizations. The answer is mostly structural, not a failure of intelligence or diligence. Patent data is, for the typical IP team, the only native dataset it owns. It arrives through tools built for patent prosecution and portfolio management, instruments designed for IP attorneys running episodic, filing-driven workflows. Those tools are genuinely excellent at the job they were built to do. They were simply never built to answer strategic, forward-looking, commercially grounded questions, because those questions were not part of the IP team's mandate when the tools were designed.
The result is a quiet optimization toward the measurable. Teams optimize for the data they can see, and white space becomes the proxy for opportunity precisely because white space is the one thing the available tooling can actually measure. Scientific momentum, commercial conviction, and forward intent are harder to see not because they are less important but because they live in datasets the IP team's tools were never wired to ingest. The gap persists because closing it has historically meant stitching together multiple disconnected platforms by hand, a manual integration burden that most teams cannot sustain quarter after quarter. So the easier path wins, and the patent map stands in for the opportunity map by default.
Closing the gap, then, is not a matter of working harder inside the patent record. No amount of additional rigor applied to a patent-only dataset produces the signals that dataset does not contain. The fix is to put the other datasets on the same surface as the patent data, so that both circles can finally be examined together rather than one at a time, and so the overlap, the actual opportunity space, becomes visible rather than inferred.
Where this is heading
The platforms built for this problem treat patents, scientific literature, and commercial signals not as separate vendor silos to be reconciled by analysts but as a single intelligence substrate. Cypris was built specifically for this, an enterprise R&D intelligence platform that unifies more than 500 million patents and scientific papers alongside commercial and market signals, grounded in a proprietary R&D ontology and serving hundreds of enterprise customers and thousands of R&D and IP professionals across Fortune 500 companies. The application most relevant to the white space problem is exactly the overlap: surfacing the gaps between heavy patent activity and heavy publication activity, and the spaces where academic or commercial momentum is building but filings have not yet appeared. Those patterns are the opportunity space, and they are invisible inside any single-source tool by construction, because no single source contains both halves of the picture.
The more recent shift is from periodic analysis toward continuous intelligence. In June 2026 Cypris launched Agentic Monitoring, which runs continuously across patent offices, scientific literature, regulatory bodies, mergers and acquisitions, product launches, grant awards, and corporate news, delivering filtered and contextualized intelligence on a defined cadence rather than waiting for a quarterly manual rebuild. The significance is not the automation in itself. It is that the strategic questions reaching the IP team do not pause between reporting cycles. Competitors hire, raise, publish, and acquire continuously, and an intelligence model that refreshes once a quarter is structurally behind the landscape it is meant to describe. Continuous monitoring closes the timing gap on the same logic that integrated data closes the coverage gap.
The role of the corporate IP team has evolved into something genuinely strategic. The mandate, the data, and the tooling are only now beginning to catch up to it. The organizations that close that gap first will be the ones making forward decisions with a forward-looking map, while their competitors are still reading the rear-view mirror and calling it the road ahead.
FAQ
What is the difference between patent white space and commercial opportunity space?
Patent white space refers to regions of a technology landscape where few or no active patents exist. Commercial opportunity space refers to areas where genuine market demand and commercial momentum are forming. The two overlap only partially, and the highest-value IP portfolios sit in the intersection where a defensible technical position meets real commercial demand. Patent data alone cannot identify that intersection because it captures only one of the two dimensions, which is why empty patent regions are routinely mistaken for open opportunities.
What is the white space fallacy?
The white space fallacy is the assumption that an empty region of the patent map represents an open commercial opportunity. An absence of patents is a starting point for investigation, not a validated opportunity. A space can be empty because there is no market, because the underlying science does not yet work, or because competitors are operating outside the patent system through trade secrets, defensive publications, or faster commercial execution. Patent data cannot distinguish between these cases, and each one demands a completely different strategic response.
Why can patent data not answer strategic R&D questions on its own?
A patent is a roughly twenty-year commitment, which makes it a forward-looking instrument, while the patent record is a backward-looking dataset that publishes filings about eighteen months after submission and reflects decisions made earlier still. Patents are lagging indicators, sometimes by years. Answering whether a technology area will still matter in five to ten years requires scientific momentum, commercial signals, and forward indicators that the patent record either omits entirely or reports too late to act on.
Has the role of the corporate IP team actually changed?
Yes, and substantially. The IP team historically protected innovations after R&D produced them, sitting downstream of the decisions that mattered. Increasingly, R&D consults IP before committing resources and treats the resulting landscape analysis as a strategic go or no-go signal. The IP function has become a strategic decision input that shapes investment direction, even though the underlying data and tooling were originally built for patent prosecution and portfolio management rather than strategy.
What datasets do IP teams need beyond patents?
Three categories. Scientific literature, including papers, preprints, grants, and clinical activity, shows where technology is heading before filings appear. Commercial signals, including venture funding, startup formation, mergers and acquisitions, and product launches, show where commercial conviction is forming. Forward indicators, including regulatory filings, clinical pipelines, market intelligence, and hiring patterns, signal intent before it becomes protected IP. Patents remain the foundation, but these three categories supply the walls the foundation alone cannot.
Why does a field with many publications but few patents matter?
A technology area with extensive recent scientific publication but limited patent filings often represents an emerging opportunity, an early window in which an organization can establish an IP position before the landscape fills in. A patent-only view registers this same area as white space and may dismiss it as empty, missing the signal entirely. The space is not empty. It is crowded with science that has not yet converted into claims.
Can hiring patterns really indicate competitive activity?
Yes, and they are among the earliest signals available. The engineering and research roles a company staffs frequently describe, in the job specifications themselves, exactly what the company is building. Because hiring precedes filing by a considerable margin, a competitor's hiring activity can reveal technical direction months or years before any of that work surfaces in the patent record.
Why does a crowded patent area still matter strategically?
A patent-dense area instinctively reads as a closed market, but contested areas are often exactly where the market is moving and where an organization may be under-protected. Density signals competitive intensity, not the absence of opportunity. Treating a crowded map as a closed door can forfeit positions where a company holds a real technical advantage but has under-filed, which can be as costly an error as treating an empty map as an open opportunity.
Why does this gap persist if it is so well understood?
The gap is structural rather than a failure of judgment. Patent data is the only native dataset most IP teams own, accessed through tools built for prosecution and portfolio management. Teams optimize for the data they can see, so white space becomes a proxy for opportunity because it is the dimension the available tooling can actually measure. Historically, closing the gap meant manually stitching together disconnected platforms quarter after quarter, a burden most teams could not sustain, so the patent-only default persisted.
How are platforms addressing the patent-only limitation?
Purpose-built R&D intelligence platforms unify patents, scientific literature, and commercial signals into a single searchable substrate rather than separate tools requiring manual reconciliation. This allows teams to see the overlap between technical defensibility and commercial momentum directly rather than inferring it. The emerging direction is continuous monitoring across patents, literature, regulatory activity, mergers and acquisitions, and corporate news, replacing periodic manual analysis with always-on intelligence that keeps pace with a landscape that never stops moving.
