Introduction to Cypris

Keep Reading

What an MCP server is, how the Model Context Protocol connectsAI assistants to patent and scientific literature databases, and how Cyprisuses MCP to deliver R&D intelligence.
AnAI assistant cannot reach live patent data on its own. Every patent searchquestion requires manual work first: pull the patent family from a database,copy the claims into the chat, ask the question, copy the answer elsewhere. TheModel Context Protocol, or MCP, removes that manual step. MCP lets an AIassistant connect directly to external data sources during a conversation.
What MCP is
MCPis an open standard for connecting AI assistants to external data sources andtools. Anthropic introduced MCP in November 2024. A data source, such as apatent database or a scientific literature index, exposes itself through an MCPserver. Any MCP-compatible AI client, including Claude Desktop and ChatGPTDesktop, can connect to that server and use it directly.
BeforeMCP, connecting an AI assistant to a specific database required a customintegration for each assistant and each data source. MCP standardizes thatconnection. One server, built once, works with any MCP-compatible client.
AnMCP server exposes three things to a connected AI client: tools it can call,such as a patent search function; resources it can read, such as patent recordsor paper abstracts; and prompts that template common tasks. A connected AIassistant can call a tool mid-conversation, retrieve current data, and answerbased on that data. It does not have to rely only on what it learned duringtraining.
Why MCP matters for patent search
Patentand scientific literature data changes constantly. A patent landscape shiftswith every new filing. A freedom-to-operate risk can appear the week before aproduct launch. A relevant paper can publish while a literature review isunderway. An AI assistant reasoning only from training data cannot know aboutany of this. It also cannot flag that its answer might be incomplete.
Patentdata is structured and authoritative. Assignee, filing date, legal status, andclaim language are facts recorded in a system of record: USPTO, EPO, WIPO.These facts do not benefit from being paraphrased from a webpage that oncementioned them. MCP lets an AI assistant query the system of record directly.It can cite exactly what it found, inside the same conversation where theanalysis is happening.
How MCP is used in patent search and R&D workflows
Aresearcher using an MCP-connected AI client can describe an invention in plainlanguage. The assistant searches live patent and literature sources directly.No query translation step is required. An IP analyst can ask about a specificassignee's recent filing activity and get an answer sourced from a current APIcall, not from training data. A scientist reviewing a technology area can pullrecent papers, patents, and citation relationships into the same conversationwhere a landscape summary is being drafted.
TheAI assistant stops operating next to the data. It starts operating on the datadirectly. Output quality depends on what data the assistant can reach throughits connected MCP server.
Where open-source MCP servers are useful, and where they stop beingenough
Open-sourceMCP servers connect AI clients to major patent and literature sources: USPTOsearch and litigation APIs, EPO's Open Patent Services for European patentdata, Google Patents, and academic sources including arXiv, PubMed, andSemantic Scholar. For a team that needs one specific data source from onespecific AI client, these are frequently the right choice. Several are activelymaintained.
Theseconnectors answer one question against one source. They do not carry contextacross a decision. A prior art search, a white space analysis, afreedom-to-operate assessment, and a regulatory check are linked stages of thesame decision: whether an R&D program is worth pursuing. The result of onestage should inform how the next is read. A single-source MCP server accuratelyreturns what its database contains. It has no framework for connecting a priorart result to a freedom-to-operate risk rating, because it answers one kind ofquery, not a workflow.
How Cypris uses MCP
Cyprisis an R&D intelligence platform, reachable through MCP, built on a corpusof more than 500 million patents and scientific papers organized through aproprietary R&D ontology. A connected AI client using Cypris through MCPworks with structured domain context, not raw results from a single searchendpoint.
Theagents available through Cypris's MCP server map to the stage-gate decisions anR&D or IP team makes: prior art review, white space identification,freedom-to-operate risk assessment, and regulatory tracking. Cypris Q, theplatform's agentic layer, and enterprise API partnerships with OpenAI,Anthropic, and Google make Cypris accessible inside the AI environmentsenterprise R&D and IP teams already use. Cypris meets enterprise-gradesecurity requirements and serves hundreds of enterprise customers acrosspharmaceuticals, chemicals, advanced materials, energy, and other regulated,security-conscious industries.
Asingle-source, open MCP server is the right tool for retrieval from one patentoffice or literature source inside one AI client. Cypris is built for adifferent need: an AI assistant that carries domain context across prior art,white space, freedom-to-operate, and regulatory decisions in the same workflow.
Setting up an MCP connection
Connectingan MCP-compatible AI client to a data source is a configuration step. Point theclient at the server. Authenticate if the source requires it. Its tools becomeavailable in conversation. Cypris is accessed through enterprise APIpartnerships rather than a self-hosted connection. This is what allows Cypristo meet enterprise security requirements while functioning as an MCP serverinside a team's existing AI client.
FAQ
**What is MCP?** MCP, theModel Context Protocol, is an open standard that lets an AI assistant connectdirectly to external data sources and tools during a conversation. Anthropicintroduced MCP in November 2024. MCP replaces custom, one-off integrations witha single protocol that works across MCP-compatible AI clients and MCP servers.
**Whatis an MCP server?** An MCP server is a connector, built on the Model ContextProtocol, that exposes a data source or tool to an MCP-compatible AI client.For patent search and R&D intelligence, an MCP server can expose patentdatabases, scientific literature indexes, or a broader intelligence platformlike Cypris to an AI assistant such as Claude Desktop or ChatGPT Desktop.
**How is MCP different froma standard API integration?** A standard integration is built once for oneapplication to connect to one data source. MCP standardizes the connection. AnyMCP-compatible AI client can use any MCP server without a new integration foreach pairing.
**Whydoes MCP matter for patent search?** Patent and scientific literature datachanges continuously. It is only useful when current and verifiable against asystem of record. An AI assistant reasoning from training data alone cannotreflect a recent filing. MCP lets the assistant query authoritative sourcesdirectly and answer based on what it retrieves.
**Does connecting to an MCPserver guarantee accurate patent search results?** No. MCP determines whetheran AI assistant can reach a data source in real time. It does not determine howcomplete that source is. A single-source MCP server accurately returns whatthat one source contains. That is not the same as complete patent landscapecoverage.
**Whatis the difference between an MCP connector and an R&D intelligence platformlike Cypris?** A connector answers one query against one data source. Cyprissupports a decision process where prior art, white space, freedom-to-operate,and regulatory findings inform each other. Cypris runs on a corpus of more than500 million patents and scientific papers organized through a proprietaryR&D ontology, delivered through an MCP server and enterprise APIpartnerships with OpenAI, Anthropic, and Google.
**Can Cypris be usedtogether with open-source MCP servers?** Yes. Teams often use open-source,single-source MCP connectors for specific databases alongside Cypris forworkflows that require reasoning across multiple linked patent and R&Ddecisions.
**Do I need to be a developer to use Cypris throughMCP?** No. Once Cypris is connected inside a compatible AI client, using it isa natural-language conversation. Cypris is accessed through enterprise APIpartnerships built to remove setup

The solid-state battery race is being decided at the electrolyte, and the patent landscape divides along three chemistries: sulfide, oxide, and polymer. A solid-state battery replaces the liquid electrolyte of a conventional lithium-ion cell with a solid one, which can improve safety and enable higher-energy electrode pairings. The central engineering problem is that no single solid electrolyte class simultaneously optimizes the three properties that matter, room-temperature ionic conductivity, stability at the electrode interfaces, and manufacturability, so each class represents a different set of trade-offs and a different region of the patent landscape. Understanding where filing activity concentrates by class, and where it does not, is how R&D and IP teams locate defensible positions in one of the fastest-moving areas of energy patenting.
The scale of that activity is documented in primary data. A joint analysis by the European Patent Office and the International Energy Agency found that international patent families in electricity storage grew from 1,029 in 2000 to more than 7,000 in 2018, at an average of 14 percent per year between 2005 and 2018, roughly four times the economy-wide average.¹ Within that, solid-state lithium-ion filings grew faster still, at around 25 percent per year since 2010, reaching 211 international patent families in 2018, with Japan the dominant country of origin, and solid-state electrolyte activity rose several-fold over the decade.¹ More recent analysis reports that energy storage now accounts for roughly 40 percent of all energy-related patenting, confirming that the field has continued to accelerate.² Because applications publish about eighteen months after filing, the most recent activity is under-represented, so these figures understate the current state.
The three electrolyte classes occupy distinct positions defined by their physics. Sulfide electrolytes reach the highest room-temperature ionic conductivities, on the order of 10 to the minus two siemens per centimeter, comparable to or exceeding liquid electrolytes, but they are chemically and electrochemically unstable at the electrode interfaces and sensitive to moisture, so the dominant patenting and research effort targets interfacial stabilization and dry-processing manufacture.³,⁴ Oxide electrolytes, principally garnet-type structures, offer good stability and a wide electrochemical window with intermediate conductivity, typically in the 10 to the minus four to 10 to the minus three siemens per centimeter range, but they are hard and brittle, which makes achieving low-resistance interfaces and scalable, thin, dense layers the central challenge.⁵ Polymer electrolytes are the most manufacturable, compatible with existing roll-to-roll processing, but historically suffered from low room-temperature conductivity, on the order of 10 to the minus seven siemens per centimeter for early systems, though engineered solid polymer electrolytes have since reached the milli-siemens-per-centimeter range, which is why manufacturability arguments increasingly favor them despite the historical conductivity gap.⁶
What the three classes trade off
Sulfide. Highest ionic conductivity, comparable to liquid electrolytes, but poor interfacial and moisture stability; patenting concentrates on interface engineering and dry manufacturing.³
Oxide. Good stability and a wide electrochemical window with intermediate conductivity, but brittleness and interfacial resistance dominate the technical and patenting effort.⁵
Polymer. Best manufacturability and compatibility with existing processes, historically limited by low room-temperature conductivity that engineered systems are now closing.⁶
Emerging classes. Halide and composite electrolytes are an active newer area that combines properties across classes, and the interfacial-engineering literature increasingly treats all classes together.⁷
How to analyze the electrolyte landscape and find white space
Scope the analysis by electrolyte class and by the property being improved, since sulfide, oxide, and polymer activity concentrate on different problems and should be assessed separately.
Aggregate to the patent-family level and attribute to organizations, so international coverage is not double-counted and activity is correctly assigned by country and assignee.
Map patents against the underlying materials research, because solid-electrolyte advances appear in scientific literature before they are patented, so literature coverage gives the earliest signal.
Identify dense and sparse regions within each class, distinguishing crowded problems, such as sulfide interface stabilization, from open white space, such as specific composite or processing approaches.
Correct for publication lag and monitor continuously, since the most recent activity is under-represented and the field moves quickly.
Where Cypris fits
Cypris runs patent landscape and white space analysis for fast-moving fields such as solid-state batteries across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by electrolyte class and by the property being improved, and normalizes organizations to canonical entities, so a team can resolve which classes and problems are crowded and which remain open as white space. Semantic search across patents and scientific literature connects filings to the underlying materials research, which matters in solid-state batteries because advances appear in the literature before they are patented. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the class-level scoping, attribution, and gap analysis, and Agentic Monitoring tracks a defined chemistry over time and flags new patents and papers as they publish, which is essential where recent activity is under-represented by publication lag. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What are the three main solid-state battery electrolyte classes?
The three main solid-state battery electrolyte classes are sulfide, oxide, and polymer. They trade off room-temperature ionic conductivity, stability at the electrode interfaces, and manufacturability, and no single class optimizes all three. Each occupies a distinct region of the patent landscape, with halide and composite electrolytes an emerging fourth area.
How do sulfide, oxide, and polymer electrolytes compare?
Sulfide electrolytes have the highest ionic conductivity, around 10 to the minus two siemens per centimeter, but poor interfacial and moisture stability. Oxide garnets offer good stability with intermediate conductivity but are brittle. Polymers are the most manufacturable but historically had low conductivity, which engineered systems are now improving.
How fast is solid-state battery patenting growing?
Solid-state battery patenting is growing quickly. Electricity-storage international patent families grew about 14 percent per year from 2005 to 2018, four times the economy-wide average, and solid-state lithium-ion filings grew around 25 percent per year since 2010. Energy storage now accounts for roughly 40 percent of all energy-related patenting.
Why does ionic conductivity differ so much between electrolyte classes?
Ionic conductivity differs between electrolyte classes because it is governed by the material's structure and ion-transport mechanism. Sulfides allow fast ion movement and reach conductivities comparable to liquids, oxides are intermediate, and polymers historically conducted far more slowly at room temperature. Engineering has narrowed the polymer gap substantially.
Which electrolyte class is winning?
No electrolyte class has decisively won, because each optimizes different properties. Sulfides lead on conductivity, oxides on stability, and polymers on manufacturability, and patenting concentrates on each class's specific weakness. The manufacturability advantage of polymers and the conductivity of sulfides are both driving heavy activity, with the outcome still open.
How do you find white space in the solid-state electrolyte landscape?
Finding white space in the solid-state electrolyte landscape means scoping by class and by the property being improved, mapping patents and scientific literature, and identifying the sparse regions within each class. Because advances appear in research first, literature coverage gives early signal. The white space is where a specific composition or processing approach is viable but few patents yet exist.
Why does solid-state battery analysis need scientific literature?
Solid-state battery analysis needs scientific literature because electrolyte and interface advances appear in materials research before they are patented, so the literature gives the earliest signal of a viable approach. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
Why does publication lag matter in the battery patent landscape?
Publication lag matters because applications publish about eighteen months after filing, so the most recent solid-state activity is under-represented in current data. In a field growing this quickly, the latest figures understate the true state. Longer-window trends and continuous monitoring are more reliable.
Who uses solid-state battery patent landscape analysis?
Solid-state battery patent landscape analysis is used by R&D, innovation, IP, and strategy teams at battery makers, automotive and energy companies, materials developers, and their partners. It informs which electrolyte class to pursue, where to file, and where competitors are concentrated. Cypris serves hundreds of enterprise customers across energy, advanced materials, chemicals, and other regulated industries.
Endnotes
- International Energy Agency & European Patent Office (2020). Innovation in Batteries and Electricity Storage: A Global Analysis Based on Patent Data. https://www.iea.org/reports/innovation-in-batteries-and-electricity-storage
- International Energy Agency (2026). The State of Energy Innovation 2026. https://www.iea.org/reports/the-state-of-energy-innovation-2026
- Richter, F. H. et al. (2020). Interfacial challenges for all-solid-state batteries based on sulfide solid electrolytes. Journal of Materiomics. https://doi.org/10.1016/j.jmat.2020.09.003
- Gamo, H., Nagai, A. & Matsuda, A. (2023). Toward Scalable Liquid-Phase Synthesis of Sulfide Solid Electrolytes for All-Solid-State Batteries. Batteries. https://doi.org/10.3390/batteries9070355
- Wei, Z. et al. (2024). Oxide Solid Electrolytes in Solid-State Batteries. Batteries & Supercaps. https://doi.org/10.1002/batt.202400667
- Wei, Z., Guo, R., Li, C. & Peng, H. (2025). Why Will Polymers Win the Race for Solid-State Batteries? Advanced Science. https://doi.org/10.1002/advs.202510481
- Chae, S. et al. (2026). Interfacial Engineering for Layered Oxide Cathodes in All-Solid-State Batteries. Batteries & Supercaps. https://doi.org/10.1002/batt.70366

Semantic search for patents retrieves documents by modeled meaning rather than by lexical overlap. A Boolean or keyword query matches on the surface form of a string: it returns patents whose text contains the specified tokens, combined through operators and often expanded with truncation, proximity, and classification filters. A semantic query matches on representation: each document and the query are encoded as high-dimensional vectors, and retrieval ranks documents by vector similarity, so patents whose claims and disclosures are conceptually related are returned even when they share no keywords with the query. This distinction is decisive in patent work because equivalent subject matter is routinely described in divergent vocabulary, and applicants draft claims with deliberately broad, idiosyncratic, or coined terminology to widen scope. When relevant patents use language the searcher did not anticipate, lexical retrieval fails to surface them, and those recall failures are the primary source of risk in prior art and freedom-to-operate analysis.
The stakes are set by scale. The World Intellectual Property Organization reported roughly 3.5 million patent applications filed worldwide in 2023,¹ and scientific literature has grown at approximately 8 to 9 percent per year across recent decades,² so the candidate space a searcher must cover exceeds what manually constructed Boolean queries can reliably span. Lexical retrieval trades recall for precision: it returns documents matching the exact terms and omits everything phrased differently. Evaluations in the patent-retrieval literature have found that keyword and Boolean logic disregard the syntax and semantics of technical language, which makes patent retrieval materially harder than general-domain information retrieval and leaves relevant documents unretrieved.³,⁴ In prior art and freedom-to-operate, the cost of a single missed document is high, because one overlooked reference can defeat a novelty position or expose a product to infringement liability.
Semantic search also underpins the shift toward agentic AI in R&D and IP. AI agents increasingly query patent and scientific corpora through an API and execute multi-step retrieval and analysis, and dense semantic retrieval is what allows an agent to locate relevant documents without a human hand-crafting Boolean strings. The Model Context Protocol (MCP), introduced by Anthropic in November 2024 and donated to the Linux Foundation's Agentic AI Foundation in December 2025, is now the common standard through which agents connect to external data.⁵ Semantic retrieval is the layer that grounds these agentic workflows in real documents.
How semantic search works
Semantic search rests on learned vector representations. A transformer-based language model, typically fine-tuned on patent and scientific text, encodes each document into a dense embedding, a fixed-length vector of several hundred dimensions whose geometry captures semantic relationships. The query is encoded by the same model, and relevance is scored by a similarity function, most commonly cosine similarity or inner product, over the shared vector space. Because encoding is done offline, retrieval at query time reduces to a nearest-neighbor search: the system returns the documents whose embeddings lie closest to the query embedding, which are the documents most similar in modeled meaning rather than in wording.
At corpus scale, exhaustive comparison against every vector is infeasible, so semantic search relies on approximate nearest-neighbor (ANN) indexing, using structures such as hierarchical navigable small-world graphs or inverted-file quantization to return near-optimal neighbors in sub-linear time. The retrieval model itself is usually a bi-encoder, which embeds query and document independently for speed. A second, more expensive cross-encoder is often applied as a reranking stage over the top candidates, jointly attending to query and document to refine the ordering. This retrieve-then-rerank architecture is standard because it combines the throughput of ANN retrieval with the precision of pairwise scoring.
Dense retrieval demonstrably outperforms lexical baselines. Dense passage retrieval improved top-20 retrieval accuracy by 9 to 19 percentage points over a strong keyword baseline in open-domain benchmarks,⁶ and patent-specific embedding models have been engineered for this domain: the European Patent Office's SEARCHFORMER uses siamese transformer encoders to produce semantic patent embeddings purpose-built for prior art search,⁷ and subsequent work has shown that few-shot fine-tuning and quantized embeddings can adapt these models efficiently to patent retrieval benchmarks.⁸ Three domain-specific factors make this adaptation necessary. Patents are long and structurally heterogeneous, so documents are segmented into passages, and claims are frequently indexed separately because claim language, not the abstract, defines legal scope. Patent corpora are multilingual, so cross-lingual embeddings allow a query in one language to retrieve prior art in another. And patents carry structured metadata, so production systems typically use hybrid retrieval, fusing dense semantic scores with lexical signals and Cooperative Patent Classification or International Patent Classification codes through rank-fusion methods to combine conceptual recall with exact-match precision.
Semantic search becomes substantially more powerful when the vector layer is combined with an explicit knowledge structure. An ontology that formalizes how technologies, materials, methods, and claims relate lets a platform interpret a query within a technology domain and cluster results by concept rather than surface wording. This combination, dense retrieval organized by an ontology, is what separates a modern patent-analytics platform from a keyword database with a search box: it improves recall by retrieving conceptually related work, and it improves interpretability by organizing results into a navigable conceptual structure. Retrieval quality in this setting is measured with recall-oriented metrics such as recall@k and mean average precision, which reflect the priority of finding all relevant documents rather than only the top few.
Where semantic search changes patent work
Prior art search. Prior art search is recall-bound, because the objective is to surface any earlier disclosure bearing on novelty or obviousness. Dense semantic retrieval surfaces prior art expressed in different terminology from the invention, including non-patent literature, which lexical search omits, producing a more complete novelty assessment.
Freedom-to-operate. Freedom-to-operate depends on identifying active, in-force claims that a product could read on, including claims drafted to cover a concept broadly. Semantic retrieval locates relevant claims independent of exact terminology and, combined with claim-level indexing, narrows the analysis to the independent claims that define infringement scope, reducing the coverage gaps that generate FTO risk.
White space analysis. White space analysis depends on clustering activity by concept to expose genuine gaps. Embedding-based clustering distinguishes real conceptual sparsity from apparent sparsity that is only an artifact of divergent terminology, so the identified white space reflects unclaimed technical territory rather than a vocabulary mismatch.
Technology and competitive intelligence. Characterizing a technology area requires connecting related work across patents and scientific literature. Encoding both sources in a shared vector space lets a platform retrieve and align conceptually related documents across them, supporting attribution of activity to technology areas and organizations.
Where Cypris fits
Cypris applies semantic search across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology is what makes the dense retrieval interpretable: it maps how technologies, claims, and research relate, so retrieval is organized by concept rather than surface wording, and results are returned as a navigable conceptual structure rather than a flat ranked list. This lets Cypris surface conceptually relevant patents and papers for prior art, freedom-to-operate at the claim level, and white space analysis, closing the recall gaps that lexical search leaves. Cypris Q, the platform's agentic layer, lets teams run semantic queries conversationally and chain them into multi-step retrieval and analysis, and Agentic Monitoring keeps results current by tracking a technology area over time. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, so AI agents can execute semantic retrieval across the corpus programmatically, and it is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is semantic search for patents?
Semantic search for patents retrieves patents and scientific papers by modeled meaning rather than by exact keywords. It encodes documents and queries as dense vectors and ranks results by vector similarity, so conceptually related documents are returned even when they share no keywords. This closes the recall gaps that cause missed prior art and freedom-to-operate risk in Boolean search.
How is semantic search different from keyword search?
Semantic search differs from keyword search in what it matches. Keyword and Boolean search match the surface form of a query string, while semantic search matches learned vector representations of meaning. For patents this matters because equivalent subject matter is described in divergent vocabulary, and lexical search misses documents phrased differently or drafted with deliberately broad claim language.
What are embeddings in semantic patent search?
Embeddings in semantic patent search are dense vectors, produced by a transformer language model, that encode the meaning of a patent, a claim, a passage, or a query into a shared high-dimensional space. Similarity between embeddings, typically cosine similarity, measures conceptual relatedness. Retrieval returns the documents whose embeddings are nearest to the query embedding.
What is dense retrieval and how does it compare to BM25?
Dense retrieval encodes queries and documents as learned vectors and ranks by vector similarity, whereas BM25 is a sparse, term-frequency lexical method. Dense passage retrieval has been shown to improve top-20 retrieval accuracy by 9 to 19 percentage points over a strong lexical baseline. Production patent systems often combine the two in hybrid retrieval to gain both conceptual recall and exact-match precision.
Why does semantic search matter for prior art search?
Semantic search matters for prior art search because prior art is recall-bound, and relevant disclosures are frequently phrased differently from the invention or appear in non-patent literature. Lexical search omits these, leaving gaps in the novelty assessment. Dense semantic retrieval surfaces conceptually related disclosures regardless of wording, making the assessment more complete.
How does semantic search improve freedom-to-operate analysis?
Semantic search improves freedom-to-operate analysis by locating active claims a product could read on even when those claims use different terminology or cover a concept broadly. Combined with claim-level indexing, it focuses the analysis on the independent claims that define infringement scope. This reduces the coverage gaps that are the main source of FTO risk.
What is a retrieve-then-rerank pipeline?
A retrieve-then-rerank pipeline is a two-stage architecture. A fast bi-encoder retrieves a candidate set using approximate nearest-neighbor search, then a more expensive cross-encoder rescores the top candidates by jointly attending to the query and each document. This combines the throughput of vector retrieval with the precision of pairwise relevance scoring.
Does semantic search need an ontology?
Semantic search does not strictly require an ontology, but combining the two is substantially more powerful. An ontology formalizes how technologies relate, letting a platform interpret a query within a domain and cluster results by concept. Cypris combines dense semantic retrieval with a proprietary R&D ontology across more than 500 million patents and scientific papers for this reason.
How do AI agents use semantic search for patents?
AI agents use semantic search as the retrieval layer that lets them locate relevant patents and papers without a human writing Boolean strings, enabling autonomous multi-step analysis. Agents query the corpus through an API and use dense retrieval to ground their reasoning in real documents. Cypris offers enterprise API partnerships with OpenAI, Anthropic, and Google so agents can run semantic retrieval across its corpus.
Which teams benefit from semantic patent search?
Semantic patent search benefits R&D, innovation, and IP teams running prior art, freedom-to-operate, white space, and technology-intelligence searches, all of which depend on high-recall retrieval of conceptually relevant documents. It is most valuable in research-intensive industries such as pharmaceuticals, chemicals, advanced materials, and energy. Cypris serves hundreds of enterprise customers across these industries.
Endnotes
- World Intellectual Property Organization. World Intellectual Property Indicators (annual series). https://www.wipo.int/publications/
- Bornmann, L. & Mutz, R. (2015). Growth rates of modern science: a bibliometric analysis based on the number of publications and cited references. Journal of the Association for Information Science and Technology. https://doi.org/10.1002/asi.23329
- Zihayat, M. & Etwaroo, R. (2021). A non-factoid question answering system for prior art search. Expert Systems with Applications. https://doi.org/10.1016/j.eswa.2021.114910
- Lupu, M., Piroi, F., Hanbury, A. & Zenz, V. (2011). CLEF-IP 2011: Retrieval in the Intellectual Property Domain.
- Anthropic (2025). Donating the Model Context Protocol and establishing the Agentic AI Foundation. https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation; Linux Foundation (2025). Formation of the Agentic AI Foundation. https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation
- Karpukhin, V. et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP. https://doi.org/10.18653/v1/2020.emnlp-main.550
- Vowinckel, K. & Hähnke, V. D. (2023). SEARCHFORMER: Semantic patent embeddings by siamese transformers for prior art search. World Patent Information. https://doi.org/10.1016/j.wpi.2023.102192
- Chikkamath, R. et al. (2025). Patent Retrieval with Few-Shot Fine-Tuning and Quantized Embeddings. https://doi.org/10.1145/3787279.3787295
