A powerful new foundation for custom queries—built on Lucene and designed for R&D precision.
Over the past few years, Cypris has helped innovation teams make faster, more informed decisions by centralizing critical insights across datasets like patents, academic papers, and company activity. But until now, our search experience relied on a legacy query system with limited capabilities, offering little support for advanced search features or dataset-level customization.
Today, we’re excited to introduce an upgraded Advanced Search on Cypris, a complete overhaul of our query engine and search experience, powered by the open-standard Lucene query syntax. This update introduces a more robust and flexible search foundation, unlocking new ways to query data, build complex filters, and extract precisely what you need across patents, research, and more.
Why we rebuilt our search system from the ground up
Cypris’ original query syntax, a proprietary format used internally for years, limited users’ ability to craft advanced queries or tailor searches to specific datasets. It lacked modern capabilities like proximity searches, field-level customization, or true Boolean logic. This made it difficult to build a reliable and intuitive experience for both casual users and advanced researchers.
By moving to Lucene, we’re adopting a powerful, industry-standard query language that makes it easier for developers to build advanced features—and gives users access to a far more capable and flexible search toolset.
What’s new in Advanced Search
1. Custom Queries by Dataset
You can now layer queries to search across datasets or tailor filters to each one. For example, you can run a broad query on drone delivery, and then add separate layers to focus on patents by a specific assignee and papers from a specific country or funding agency.
Navigating the All Datasets tab introduces a new level of complexity—and power—by allowing users to apply dataset-specific logic within a single, unified query workflow. While querying multiple datasets simultaneously might seem straightforward, the underlying differences in schema, metadata, and available fields between our proprietary datasets make this a deeply technical challenge. Patents, for example, include claims, application numbers, and multiple date fields (filed, granted, updated), while academic papers use DOIs, have different structural conventions, and emphasize different metadata. In the past, we sidestepped this complexity by translating general queries like ((drone_allText)) into dataset-specific logic under the hood. Now, instead of obscuring that logic, we allow users to opt in to it. The builder provides progressive layers of customization: start with intuitive keyword searches across all fields, then move into the advanced builder for field-specific targeting, fuzzy logic, and term boosting, and finally, tailor query logic by dataset—such as specifying different countries of interest for papers vs. patents. This approach preserves flexibility while giving users full control, and with tools like our real-time Live Analysis and “Your Query” panel, we make it easy to understand how every decision affects the results.
2. More Fields to Query
We’re exposing deeper fields across datasets—giving you explicit control over the dimensions of your search. For the first time, users can now search academic papers by DOI, a critical identifier previously unsupported on the platform. You can also query by:
- Author or inventor names
- Organizations or assignees
- Countries, journals, funding agencies, and more
3. Full Boolean Support
Advanced Search now leverages powerful Boolean logic—AND, OR, NOT, and grouping—enabling more precise control over search logic and improving performance and accuracy.
4. Lucene Syntax Features
Use built-in Lucene features to create expressive, complex searches:
- Proximity searches to find terms near each other
- Fuzzy searches for flexible matching
- Exact phrase matching
- Boosting to prioritize results (e.g., prioritize results mentioning AI 3x more than others)
- Prefix/Postfix queries to match phrases that start or end a certain way
- Range queries for fields like date, funding amounts, or numerical values
A more powerful user experience
Our new search interface is built to help you tap into these capabilities without needing to know the syntax from the start. You’ll find:
- A Query Builder to guide you through complex searches
- A Help Video to onboard users to Lucene-style searches
- Inline examples and tips for writing queries using grouping, boosting, and more
Built for precision, speed, and customization
With Lucene as our foundation, search results are now not only more flexible but also faster and more accurate. Semantic search continues to offer natural-language ease of use, while Boolean search gives power users the performance and structure they need to uncover insights with greater specificity.
Whether you’re an innovation analyst drilling into AI patents or a business development lead scanning academic papers from Chilean researchers—Advanced Search is built to help you get to the signal, faster.
Available now to all users
Advanced Search is live and available across the Cypris platform today. If you’re already using Cypris, you’ll find the new search interface in your dashboard, complete with updated syntax documentation and walkthroughs.
We’re excited to see what you’ll build, discover, and analyze with this new capability. This is just the beginning—we’ll continue expanding the fields, syntax features, and customization options as we push the boundaries of what intelligent search can do for R&D.

Introducing Advanced Search on Cypris

A powerful new foundation for custom queries—built on Lucene and designed for R&D precision.
Over the past few years, Cypris has helped innovation teams make faster, more informed decisions by centralizing critical insights across datasets like patents, academic papers, and company activity. But until now, our search experience relied on a legacy query system with limited capabilities, offering little support for advanced search features or dataset-level customization.
Today, we’re excited to introduce an upgraded Advanced Search on Cypris, a complete overhaul of our query engine and search experience, powered by the open-standard Lucene query syntax. This update introduces a more robust and flexible search foundation, unlocking new ways to query data, build complex filters, and extract precisely what you need across patents, research, and more.
Why we rebuilt our search system from the ground up
Cypris’ original query syntax, a proprietary format used internally for years, limited users’ ability to craft advanced queries or tailor searches to specific datasets. It lacked modern capabilities like proximity searches, field-level customization, or true Boolean logic. This made it difficult to build a reliable and intuitive experience for both casual users and advanced researchers.
By moving to Lucene, we’re adopting a powerful, industry-standard query language that makes it easier for developers to build advanced features—and gives users access to a far more capable and flexible search toolset.
What’s new in Advanced Search
1. Custom Queries by Dataset
You can now layer queries to search across datasets or tailor filters to each one. For example, you can run a broad query on drone delivery, and then add separate layers to focus on patents by a specific assignee and papers from a specific country or funding agency.
Navigating the All Datasets tab introduces a new level of complexity—and power—by allowing users to apply dataset-specific logic within a single, unified query workflow. While querying multiple datasets simultaneously might seem straightforward, the underlying differences in schema, metadata, and available fields between our proprietary datasets make this a deeply technical challenge. Patents, for example, include claims, application numbers, and multiple date fields (filed, granted, updated), while academic papers use DOIs, have different structural conventions, and emphasize different metadata. In the past, we sidestepped this complexity by translating general queries like ((drone_allText)) into dataset-specific logic under the hood. Now, instead of obscuring that logic, we allow users to opt in to it. The builder provides progressive layers of customization: start with intuitive keyword searches across all fields, then move into the advanced builder for field-specific targeting, fuzzy logic, and term boosting, and finally, tailor query logic by dataset—such as specifying different countries of interest for papers vs. patents. This approach preserves flexibility while giving users full control, and with tools like our real-time Live Analysis and “Your Query” panel, we make it easy to understand how every decision affects the results.
2. More Fields to Query
We’re exposing deeper fields across datasets—giving you explicit control over the dimensions of your search. For the first time, users can now search academic papers by DOI, a critical identifier previously unsupported on the platform. You can also query by:
- Author or inventor names
- Organizations or assignees
- Countries, journals, funding agencies, and more
3. Full Boolean Support
Advanced Search now leverages powerful Boolean logic—AND, OR, NOT, and grouping—enabling more precise control over search logic and improving performance and accuracy.
4. Lucene Syntax Features
Use built-in Lucene features to create expressive, complex searches:
- Proximity searches to find terms near each other
- Fuzzy searches for flexible matching
- Exact phrase matching
- Boosting to prioritize results (e.g., prioritize results mentioning AI 3x more than others)
- Prefix/Postfix queries to match phrases that start or end a certain way
- Range queries for fields like date, funding amounts, or numerical values
A more powerful user experience
Our new search interface is built to help you tap into these capabilities without needing to know the syntax from the start. You’ll find:
- A Query Builder to guide you through complex searches
- A Help Video to onboard users to Lucene-style searches
- Inline examples and tips for writing queries using grouping, boosting, and more
Built for precision, speed, and customization
With Lucene as our foundation, search results are now not only more flexible but also faster and more accurate. Semantic search continues to offer natural-language ease of use, while Boolean search gives power users the performance and structure they need to uncover insights with greater specificity.
Whether you’re an innovation analyst drilling into AI patents or a business development lead scanning academic papers from Chilean researchers—Advanced Search is built to help you get to the signal, faster.
Available now to all users
Advanced Search is live and available across the Cypris platform today. If you’re already using Cypris, you’ll find the new search interface in your dashboard, complete with updated syntax documentation and walkthroughs.
We’re excited to see what you’ll build, discover, and analyze with this new capability. This is just the beginning—we’ll continue expanding the fields, syntax features, and customization options as we push the boundaries of what intelligent search can do for R&D.

Keep Reading

Semantic search for patents retrieves documents by modeled meaning rather than by lexical overlap. A Boolean or keyword query matches on the surface form of a string: it returns patents whose text contains the specified tokens, combined through operators and often expanded with truncation, proximity, and classification filters. A semantic query matches on representation: each document and the query are encoded as high-dimensional vectors, and retrieval ranks documents by vector similarity, so patents whose claims and disclosures are conceptually related are returned even when they share no keywords with the query. This distinction is decisive in patent work because equivalent subject matter is routinely described in divergent vocabulary, and applicants draft claims with deliberately broad, idiosyncratic, or coined terminology to widen scope. When relevant patents use language the searcher did not anticipate, lexical retrieval fails to surface them, and those recall failures are the primary source of risk in prior art and freedom-to-operate analysis.
The stakes are set by scale. The World Intellectual Property Organization reported roughly 3.5 million patent applications filed worldwide in 2023,¹ and scientific literature has grown at approximately 8 to 9 percent per year across recent decades,² so the candidate space a searcher must cover exceeds what manually constructed Boolean queries can reliably span. Lexical retrieval trades recall for precision: it returns documents matching the exact terms and omits everything phrased differently. Evaluations in the patent-retrieval literature have found that keyword and Boolean logic disregard the syntax and semantics of technical language, which makes patent retrieval materially harder than general-domain information retrieval and leaves relevant documents unretrieved.³,⁴ In prior art and freedom-to-operate, the cost of a single missed document is high, because one overlooked reference can defeat a novelty position or expose a product to infringement liability.
Semantic search also underpins the shift toward agentic AI in R&D and IP. AI agents increasingly query patent and scientific corpora through an API and execute multi-step retrieval and analysis, and dense semantic retrieval is what allows an agent to locate relevant documents without a human hand-crafting Boolean strings. The Model Context Protocol (MCP), introduced by Anthropic in November 2024 and donated to the Linux Foundation's Agentic AI Foundation in December 2025, is now the common standard through which agents connect to external data.⁵ Semantic retrieval is the layer that grounds these agentic workflows in real documents.
How semantic search works
Semantic search rests on learned vector representations. A transformer-based language model, typically fine-tuned on patent and scientific text, encodes each document into a dense embedding, a fixed-length vector of several hundred dimensions whose geometry captures semantic relationships. The query is encoded by the same model, and relevance is scored by a similarity function, most commonly cosine similarity or inner product, over the shared vector space. Because encoding is done offline, retrieval at query time reduces to a nearest-neighbor search: the system returns the documents whose embeddings lie closest to the query embedding, which are the documents most similar in modeled meaning rather than in wording.
At corpus scale, exhaustive comparison against every vector is infeasible, so semantic search relies on approximate nearest-neighbor (ANN) indexing, using structures such as hierarchical navigable small-world graphs or inverted-file quantization to return near-optimal neighbors in sub-linear time. The retrieval model itself is usually a bi-encoder, which embeds query and document independently for speed. A second, more expensive cross-encoder is often applied as a reranking stage over the top candidates, jointly attending to query and document to refine the ordering. This retrieve-then-rerank architecture is standard because it combines the throughput of ANN retrieval with the precision of pairwise scoring.
Dense retrieval demonstrably outperforms lexical baselines. Dense passage retrieval improved top-20 retrieval accuracy by 9 to 19 percentage points over a strong keyword baseline in open-domain benchmarks,⁶ and patent-specific embedding models have been engineered for this domain: the European Patent Office's SEARCHFORMER uses siamese transformer encoders to produce semantic patent embeddings purpose-built for prior art search,⁷ and subsequent work has shown that few-shot fine-tuning and quantized embeddings can adapt these models efficiently to patent retrieval benchmarks.⁸ Three domain-specific factors make this adaptation necessary. Patents are long and structurally heterogeneous, so documents are segmented into passages, and claims are frequently indexed separately because claim language, not the abstract, defines legal scope. Patent corpora are multilingual, so cross-lingual embeddings allow a query in one language to retrieve prior art in another. And patents carry structured metadata, so production systems typically use hybrid retrieval, fusing dense semantic scores with lexical signals and Cooperative Patent Classification or International Patent Classification codes through rank-fusion methods to combine conceptual recall with exact-match precision.
Semantic search becomes substantially more powerful when the vector layer is combined with an explicit knowledge structure. An ontology that formalizes how technologies, materials, methods, and claims relate lets a platform interpret a query within a technology domain and cluster results by concept rather than surface wording. This combination, dense retrieval organized by an ontology, is what separates a modern patent-analytics platform from a keyword database with a search box: it improves recall by retrieving conceptually related work, and it improves interpretability by organizing results into a navigable conceptual structure. Retrieval quality in this setting is measured with recall-oriented metrics such as recall@k and mean average precision, which reflect the priority of finding all relevant documents rather than only the top few.
Where semantic search changes patent work
Prior art search. Prior art search is recall-bound, because the objective is to surface any earlier disclosure bearing on novelty or obviousness. Dense semantic retrieval surfaces prior art expressed in different terminology from the invention, including non-patent literature, which lexical search omits, producing a more complete novelty assessment.
Freedom-to-operate. Freedom-to-operate depends on identifying active, in-force claims that a product could read on, including claims drafted to cover a concept broadly. Semantic retrieval locates relevant claims independent of exact terminology and, combined with claim-level indexing, narrows the analysis to the independent claims that define infringement scope, reducing the coverage gaps that generate FTO risk.
White space analysis. White space analysis depends on clustering activity by concept to expose genuine gaps. Embedding-based clustering distinguishes real conceptual sparsity from apparent sparsity that is only an artifact of divergent terminology, so the identified white space reflects unclaimed technical territory rather than a vocabulary mismatch.
Technology and competitive intelligence. Characterizing a technology area requires connecting related work across patents and scientific literature. Encoding both sources in a shared vector space lets a platform retrieve and align conceptually related documents across them, supporting attribution of activity to technology areas and organizations.
Where Cypris fits
Cypris applies semantic search across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology is what makes the dense retrieval interpretable: it maps how technologies, claims, and research relate, so retrieval is organized by concept rather than surface wording, and results are returned as a navigable conceptual structure rather than a flat ranked list. This lets Cypris surface conceptually relevant patents and papers for prior art, freedom-to-operate at the claim level, and white space analysis, closing the recall gaps that lexical search leaves. Cypris Q, the platform's agentic layer, lets teams run semantic queries conversationally and chain them into multi-step retrieval and analysis, and Agentic Monitoring keeps results current by tracking a technology area over time. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, so AI agents can execute semantic retrieval across the corpus programmatically, and it is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is semantic search for patents?
Semantic search for patents retrieves patents and scientific papers by modeled meaning rather than by exact keywords. It encodes documents and queries as dense vectors and ranks results by vector similarity, so conceptually related documents are returned even when they share no keywords. This closes the recall gaps that cause missed prior art and freedom-to-operate risk in Boolean search.
How is semantic search different from keyword search?
Semantic search differs from keyword search in what it matches. Keyword and Boolean search match the surface form of a query string, while semantic search matches learned vector representations of meaning. For patents this matters because equivalent subject matter is described in divergent vocabulary, and lexical search misses documents phrased differently or drafted with deliberately broad claim language.
What are embeddings in semantic patent search?
Embeddings in semantic patent search are dense vectors, produced by a transformer language model, that encode the meaning of a patent, a claim, a passage, or a query into a shared high-dimensional space. Similarity between embeddings, typically cosine similarity, measures conceptual relatedness. Retrieval returns the documents whose embeddings are nearest to the query embedding.
What is dense retrieval and how does it compare to BM25?
Dense retrieval encodes queries and documents as learned vectors and ranks by vector similarity, whereas BM25 is a sparse, term-frequency lexical method. Dense passage retrieval has been shown to improve top-20 retrieval accuracy by 9 to 19 percentage points over a strong lexical baseline. Production patent systems often combine the two in hybrid retrieval to gain both conceptual recall and exact-match precision.
Why does semantic search matter for prior art search?
Semantic search matters for prior art search because prior art is recall-bound, and relevant disclosures are frequently phrased differently from the invention or appear in non-patent literature. Lexical search omits these, leaving gaps in the novelty assessment. Dense semantic retrieval surfaces conceptually related disclosures regardless of wording, making the assessment more complete.
How does semantic search improve freedom-to-operate analysis?
Semantic search improves freedom-to-operate analysis by locating active claims a product could read on even when those claims use different terminology or cover a concept broadly. Combined with claim-level indexing, it focuses the analysis on the independent claims that define infringement scope. This reduces the coverage gaps that are the main source of FTO risk.
What is a retrieve-then-rerank pipeline?
A retrieve-then-rerank pipeline is a two-stage architecture. A fast bi-encoder retrieves a candidate set using approximate nearest-neighbor search, then a more expensive cross-encoder rescores the top candidates by jointly attending to the query and each document. This combines the throughput of vector retrieval with the precision of pairwise relevance scoring.
Does semantic search need an ontology?
Semantic search does not strictly require an ontology, but combining the two is substantially more powerful. An ontology formalizes how technologies relate, letting a platform interpret a query within a domain and cluster results by concept. Cypris combines dense semantic retrieval with a proprietary R&D ontology across more than 500 million patents and scientific papers for this reason.
How do AI agents use semantic search for patents?
AI agents use semantic search as the retrieval layer that lets them locate relevant patents and papers without a human writing Boolean strings, enabling autonomous multi-step analysis. Agents query the corpus through an API and use dense retrieval to ground their reasoning in real documents. Cypris offers enterprise API partnerships with OpenAI, Anthropic, and Google so agents can run semantic retrieval across its corpus.
Which teams benefit from semantic patent search?
Semantic patent search benefits R&D, innovation, and IP teams running prior art, freedom-to-operate, white space, and technology-intelligence searches, all of which depend on high-recall retrieval of conceptually relevant documents. It is most valuable in research-intensive industries such as pharmaceuticals, chemicals, advanced materials, and energy. Cypris serves hundreds of enterprise customers across these industries.
Endnotes
- World Intellectual Property Organization. World Intellectual Property Indicators (annual series). https://www.wipo.int/publications/
- Bornmann, L. & Mutz, R. (2015). Growth rates of modern science: a bibliometric analysis based on the number of publications and cited references. Journal of the Association for Information Science and Technology. https://doi.org/10.1002/asi.23329
- Zihayat, M. & Etwaroo, R. (2021). A non-factoid question answering system for prior art search. Expert Systems with Applications. https://doi.org/10.1016/j.eswa.2021.114910
- Lupu, M., Piroi, F., Hanbury, A. & Zenz, V. (2011). CLEF-IP 2011: Retrieval in the Intellectual Property Domain.
- Anthropic (2025). Donating the Model Context Protocol and establishing the Agentic AI Foundation. https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation; Linux Foundation (2025). Formation of the Agentic AI Foundation. https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation
- Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP. https://doi.org/10.18653/v1/2020.emnlp-main.550
- Vowinckel, K. & Hähnke, V. D. (2023). SEARCHFORMER: Semantic patent embeddings by siamese transformers for prior art search. World Patent Information. https://doi.org/10.1016/j.wpi.2023.102192
- Chikkamath, R. et al. (2025). Patent Retrieval with Few-Shot Fine-Tuning and Quantized Embeddings. https://doi.org/10.1145/3787279.3787295

An ontology is a formal, machine-readable specification of the concepts in a domain and the relationships among them. The term has a precise meaning in knowledge representation: an explicit specification of a conceptualization,¹ that is, a defined vocabulary of entity types, attributes, and relations, together with constraints on how they may be combined. This distinguishes an ontology from a flat taxonomy, which only arranges terms hierarchically; an ontology also encodes non-hierarchical relations, such as a material being used in a process or a method being applied to a claim. In R&D and patent intelligence, the ontology defines the domain schema: the technologies, materials, methods, claims, organizations, and research areas that matter, and the relationship types that connect them.²
A knowledge graph instantiates that schema over real data. It represents information as a graph of nodes and typed edges, commonly expressed as subject-predicate-object triples, linking specific patents, scientific papers, assignees, inventors, technologies, and materials as connected entities rather than isolated documents. Building the graph requires several engineering steps that determine its quality: named-entity recognition and relation extraction to convert unstructured patent and paper text into triples; entity resolution to normalize the many surface forms of an organization, inventor, or compound to a single canonical node; and provenance tracking so every assertion in the graph traces back to the source document that supports it. The result is a structure that can be queried declaratively, for example with a graph query language, and that supports multi-hop traversal, so a question can follow chains of relationships rather than matching a single string.
This structure matters because patents and scientific literature become intelligence only when their relationships are made explicit. A ranked list of relevant documents does not state how a technology area is organized, which organizations are active, how research connects to patents, or where the graph is sparse. An ontology-backed knowledge graph makes those relationships first-class and queryable. A team can ask how two technologies relate, which body of research underpins a patent cluster, which assignees co-file in an area, or where a domain is unclaimed, and receive an answer computed over structured connections rather than assembled by reading.
The 2026 relevance is that structured knowledge is the most reliable way to ground generative AI. Large language models produce fluent output but can assert unsupported claims when they generate from parametric memory over unstructured text. Retrieval-augmented generation (RAG), which conditions a model's output on retrieved external evidence, was introduced to address this and improves factual accuracy on knowledge-intensive tasks.³,⁴ Graph retrieval-augmented generation (GraphRAG) extends RAG by retrieving connected subgraphs rather than isolated passages, so the model reasons over entities and their relationships and can answer questions that require traversing multiple hops.⁵,⁶ Grounding a system on an ontology-backed knowledge graph constrains its outputs to real, connected entities, which is essential for patent and R&D work where every conclusion must trace to actual patents and papers, and where retrieval quality directly governs the reliability of downstream generation.⁷ It is also what makes agentic workflows dependable: an agent reasoning over a structured, provenance-tracked graph produces results a team can verify against sources.
What an ontology and knowledge graph add to patent intelligence
Multi-hop reasoning over relationships. A knowledge graph answers relational and multi-hop questions, such as how two technologies connect through shared materials or which research a patent cluster builds on, rather than only returning documents that match a query string.
Concept-organized semantic search. Dense semantic retrieval returns conceptually relevant documents; the ontology organizes that retrieval within a domain schema, improving both recall and the interpretability of results by grouping them under defined concepts.
White space analysis. White space analysis depends on clustering activity by concept to expose genuine gaps. Clustering patents and papers over the ontology's relationship structure exposes real conceptual sparsity rather than gaps that are artifacts of divergent terminology.
Entity-resolved attribution and competitive intelligence. Entity resolution normalizes assignee and inventor variants to canonical nodes, which lets the graph attribute filings and research accurately and build co-assignee and citation networks rather than a document list.
Provenance-grounded AI. The ontology and knowledge graph give AI agents a structured, provenance-tracked foundation to reason over, which improves the accuracy of agentic analysis and makes its results traceable to the specific patents and papers that support them.
Where Cypris fits
Cypris organizes a corpus of more than 500 million patents and scientific papers through a proprietary R&D ontology. That ontology is the core of the platform: it defines how technologies, claims, materials, methods, and research relate, so Cypris reasons over an entity-resolved relationship structure rather than only matching keywords. This structure powers dense semantic retrieval organized by concept, white space analysis that exposes genuine conceptual gaps, and competitive intelligence that attributes activity to canonical organizations and technology areas. Cypris Q, the platform's agentic layer, reasons over this provenance-tracked foundation, which is what makes its multi-step analysis both reliable and traceable to real patents and papers, consistent with graph-grounded retrieval approaches. Agentic Monitoring tracks a technology area over time against the same structure. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, so AI agents can query the structured corpus programmatically, and it is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is an ontology in R&D and patent intelligence?
An ontology in R&D and patent intelligence is a formal, machine-readable specification of the concepts in the domain and the relationships among them, defined as an explicit specification of a conceptualization. It sets out the entity types, such as technologies, materials, methods, and claims, and the relations that connect them. This lets a platform reason over connections between patents and scientific literature rather than treating documents as isolated.
How is an ontology different from a taxonomy?
An ontology differs from a taxonomy in expressiveness. A taxonomy arranges terms in a hierarchy, while an ontology also encodes non-hierarchical, typed relationships and constraints, such as a material being used in a process. This richer structure is what allows multi-hop reasoning across patents and research rather than simple category lookup.
What is a knowledge graph for patents?
A knowledge graph for patents represents patents, scientific papers, assignees, inventors, technologies, and materials as nodes connected by typed edges, commonly expressed as subject-predicate-object triples. It applies an ontology's schema to real data so relationships are explicit and queryable. This turns a document collection into a structure that supports declarative queries and multi-hop traversal.
How is a knowledge graph built from patent text?
A knowledge graph is built from patent text through named-entity recognition and relation extraction to convert unstructured text into triples, entity resolution to normalize variant names to canonical nodes, and provenance tracking so each assertion links back to its source document. The quality of these steps determines the reliability of the graph. Poor entity resolution, for example, fragments an organization across many nodes and distorts attribution.
Why do knowledge graphs matter for AI in patent research?
Knowledge graphs matter for AI in patent research because they ground generative models on real, connected entities, which improves accuracy and traceability. A model generating from unstructured text alone can assert unsupported claims, whereas one conditioned on a provenance-tracked graph constrains its answers to actual patents and papers. This is essential where conclusions must be verifiable.
What is GraphRAG and how does it differ from standard RAG? GraphRAG is graph retrieval-augmented generation. Standard RAG retrieves isolated text passages to condition a model's output, while GraphRAG retrieves connected subgraphs, so the model reasons over entities and their relationships and can answer multi-hop questions. This suits patent intelligence, where questions often require traversing links between technologies, research, and organizations.
How does an ontology improve white space analysis?
An ontology improves white space analysis by clustering patents and papers over defined relationships rather than by exact keywords, which exposes genuine conceptual gaps instead of gaps that are only artifacts of differing terminology. Because the sparsity reflects the domain structure, the identified white space corresponds to unclaimed technical territory. Cypris organizes its corpus of more than 500 million patents and scientific papers through a proprietary R&D ontology for this purpose.
How do knowledge graphs reduce AI hallucination in patent work?
Knowledge graphs reduce AI hallucination in patent work by constraining a model's outputs to real, connected entities with tracked provenance rather than letting it generate from unstructured text. Retrieval-augmented approaches, and graph-based retrieval in particular, condition generation on retrieved evidence, which improves factual accuracy and lets conclusions be traced to sources. This makes results verifiable against the underlying patents and papers.
Is a knowledge graph the same as a vector database?
A knowledge graph is not the same as a vector database. A vector database supports semantic similarity search over embeddings, while a knowledge graph represents explicit, typed relationships between entities. They are complementary: dense retrieval finds relevant documents, and the graph structures how those documents and entities relate. Cypris combines semantic retrieval with a proprietary R&D ontology.
Which teams benefit from ontology-based patent intelligence?
Ontology-based patent intelligence benefits R&D, innovation, IP, and strategy teams that need to understand how technologies relate, attribute activity to organizations, and find genuine white space. It is most valuable in research-intensive industries such as pharmaceuticals, chemicals, advanced materials, and energy. Cypris serves hundreds of enterprise customers across these industries.
Endnotes
- Gruber, T. R. (1993). A translation approach to portable ontology specifications. Knowledge Acquisition. https://doi.org/10.1006/knac.1993.1008
- Gruber, T. R. (1995). Toward principles for the design of ontologies used for knowledge sharing. International Journal of Human-Computer Studies. https://doi.org/10.1006/ijhc.1995.1081
- Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS.
- Gao, Y. et al. (2023). Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997. https://doi.org/10.48550/arxiv.2312.10997
- Procko, T. & Ochoa, O. (2024). Graph Retrieval-Augmented Generation for Large Language Models: A Survey. https://doi.org/10.1109/aixset62544.2024.00030
- Han, S. et al. (2025). A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models. arXiv:2501.13958. https://doi.org/10.48550/arxiv.2501.13958
- Chen, J. et al. (2024). Benchmarking Large Language Models in Retrieval-Augmented Generation. AAAI. https://doi.org/10.1609/aaai.v38i16.29728

Patent filings are a leading indicator of competitor R&D direction, and the lead time is a structural consequence of how the patent system operates. An application is filed at its priority date, well before the corresponding product reaches the market, and under the standard 18-month publication rule reflected in USPTO practice and PCT Article 21,⁵ it is not published until roughly eighteen months after that priority date. The interval between when a competitor commits R&D and when the public can observe it is therefore built into the system. The International Energy Agency treats patenting as a leading indicator of technological change in its innovation analysis,¹ and the same logic holds across sectors: a competitor's published filings reveal committed R&D direction ahead of the market, and studies of the linkage between scientific publication and patenting document a measurable lag between the two that compounds the observable lead time.² For R&D and competitive intelligence teams, this makes patents one of the most reliable forward-looking competitive signals available.
Reading that signal well requires structured analysis rather than filing counts, and several technical steps determine its accuracy. First, the unit of analysis should be the patent family, not the individual document, because a single invention generates multiple applications across jurisdictions; counting documents rather than families overstates activity and double-counts international coverage. Second, filings must be located in the technology space using classification codes, principally the Cooperative Patent Classification and International Patent Classification systems, which assign standardized technology categories independent of the applicant's terminology. Third, activity must be attributed through assignee disambiguation, normalizing the many name variants, subsidiaries, and transliterations of an organization to a single canonical entity, because unresolved assignee names fragment a competitor's portfolio and distort the picture. Fourth, the analysis should read the trend over time rather than the latest counts, because the most recent eighteen-to-twenty-four months of data are systematically under-represented by publication lag, so apparent recent declines are usually artifacts rather than real slowdowns.
Two network structures add depth beyond volume. Forward and backward citation analysis situates a competitor's filings in the flow of prior art: backward citations reveal the foundations a filing builds on, and forward citations indicate influence and where a technology is being extended. Co-assignee and knowledge-search network analysis reveals partnerships, academic-industry pipelines, and the coupling between organizations, which shape a competitor's future direction; network-embedding methods over these structures are an established competitive-intelligence technique.³ Scientific literature strengthens the signal further, because research is published before it is patented and patents are filed before products ship, so combining the two sources extends the observable lead time; the scientific footprint within a competitor's filings can be traced through their non-patent references.⁴
What competitor filings reveal
Technology direction. The classification areas where a competitor is filing show where R&D is being committed, often well before those commitments appear in products.
Intensity and momentum. The distribution and rate of change of filing activity across technology areas indicate priorities, and shifts in filing momentum signal changes in strategy earlier than raw counts.
Adjacent moves. Filings in classifications adjacent to a competitor's current products can signal diversification or expansion before it is announced.
Research foundations. The non-patent references and scientific literature a competitor's filings build on show the research base behind their direction, and rising related research is an earlier signal still.
Collaboration structure. Co-assignee patterns and citation coupling reveal partnerships and academic-industry pipelines; network analysis of these relationships is an established competitive-intelligence method.³
How to read competitor R&D direction
Define the competitors and the technology space, scoping the latter with classification codes so the boundary is standardized and reproducible.
Resolve assignees to canonical entities and aggregate to the patent-family level, so activity is attributed accurately and international coverage is not double-counted.
Cluster filings by concept using semantic analysis over the classification and text, so related work groups together regardless of terminology.
Analyze filing momentum as a time series, discounting the most recent windows for publication lag, since direction is visible in trends rather than in the latest bar.
Connect filings to their non-patent references and to the scientific literature, to extend the lead time and expose the research foundations.
Monitor continuously, because competitor direction is revealed by how activity shifts, and continuous monitoring captures those shifts as they publish.
Where Cypris fits
Cypris supports competitive intelligence across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology and its entity resolution are what turn filings into direction: they normalize assignees to canonical organizations, aggregate to the family level, and cluster activity by concept, so a team sees where a competitor is moving rather than a list of documents. Dense semantic search across patents and scientific literature connects filings to their research foundations, which extends the lead time on the signal, and citation and co-assignee structures expose collaboration and influence. Cypris Q, the platform's agentic layer, lets teams analyze competitor direction conversationally and chain the attribution, clustering, and time-series analysis. Agentic Monitoring is central to this use case: it tracks defined competitors and technology areas over time and flags new filings and research as they publish, so competitive intelligence is continuous rather than a one-time report. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, so AI agents can query the corpus programmatically, and it is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
How do patent filings reveal competitor R&D direction?
Patent filings reveal competitor R&D direction because an application is filed at its priority date, before the product ships, and is published only about eighteen months later under the standard publication rule. This built-in lag means published filings show committed R&D ahead of the market. Reading the direction requires attributing filings to competitors and technology areas and analyzing where activity concentrates and shifts.
What is the 18-month publication rule?
The 18-month publication rule is the standard practice, reflected in USPTO procedure and PCT Article 21, under which a patent application is published approximately eighteen months after its earliest priority date. It creates a predictable interval between filing and public visibility. It is also why the most recent windows of filing data are under-represented and should not be read as slowdowns.
Why analyze patent families instead of individual documents?
Analyzing patent families instead of individual documents avoids double-counting, because a single invention generates multiple applications across jurisdictions. Counting documents overstates activity and conflates international coverage with genuine volume. The family is the correct unit for measuring how much distinct R&D a competitor is committing.
What role do classification codes play?
Classification codes, principally the Cooperative Patent Classification and International Patent Classification systems, assign standardized technology categories to filings independent of the applicant's wording. They let an analyst locate and compare activity in a technology space reproducibly. This is more reliable than keyword filtering, which varies with drafting style.
Why is assignee disambiguation important?
Assignee disambiguation is important because organizations appear under many name variants, subsidiaries, and transliterations, and unresolved names fragment a competitor's portfolio across multiple entities. Normalizing these to a single canonical entity is what makes attribution and trend analysis accurate. Poor disambiguation systematically distorts competitive intelligence.
How do citation networks support competitive intelligence?
Citation networks support competitive intelligence by situating filings in the flow of prior art. Backward citations reveal the foundations a filing builds on, and forward citations indicate influence and where a technology is being extended. Co-assignee and knowledge-search network analysis additionally reveals partnerships and academic-industry pipelines.
Why combine patents with scientific literature?
Combining patents with scientific literature extends the observable lead time, because research is published before it is patented and patents precede products. Rising research associated with a competitor, followed by early filings, is an earlier and stronger signal than filings alone. The scientific footprint within filings can be traced through their non-patent references.
Why not just count competitor patent filings?
Counting filings alone is misleading because recent counts are depressed by publication lag and raw volume does not indicate direction. The informative signal is which classification areas activity concentrates in and how that distribution changes over time. Family-level aggregation, classification analysis, and time-series momentum are what reveal direction.
Why is continuous monitoring important for competitive intelligence? Continuous monitoring is important because competitor direction is revealed by how activity changes, which a one-time report cannot capture, and because new filings and research publish constantly. A shift in a competitor's focus is only visible if the area is tracked over time. Cypris uses Agentic Monitoring to track competitors and technology areas and flag new activity as it publishes.
Which teams read competitor R&D direction from patents?
Reading competitor R&D direction from patents is done by competitive intelligence, R&D, innovation, strategy, and corporate development teams that need forward-looking awareness of competitor moves. It is most valuable in research-intensive industries such as pharmaceuticals, chemicals, advanced materials, and energy. Cypris serves hundreds of enterprise customers across these industries.
Endnotes
- International Energy Agency (2026). The State of Energy Innovation 2026. https://www.iea.org/reports/the-state-of-energy-innovation-2026
- Fukuzawa, N. & Ida, T. (2015). Science linkages between scientific articles and patents for leading scientists in the life and medical sciences field. Scientometrics. https://doi.org/10.1007/s11192-015-1795-z
- Yang, X. et al. (2024). Predicting patent transaction behaviour based on embedded features of knowledge search networks. Journal of Knowledge Management. https://doi.org/10.1108/jkm-12-2023-1220
- Callaert, J., Grouwels, J. & Van Looy, B. (2011). Delineating the scientific footprint in technology: identifying scientific publications within non-patent references. Scientometrics. https://doi.org/10.1007/s11192-011-0573-9
- World Intellectual Property Organization, PCT Article 21 (International Publication), and USPTO Manual of Patent Examining Procedure, on patent publication timing.
