# OriginChainDB > OriginChainDB is a distributed multimodal database for connected applications and production AI. SQL, vector similarity, graph traversal, and full-text search share one managed engine. Natural-language `/ask` provides another way to query that data. OriginChainDB is operated by Silicoyn Technologies Pvt Ltd. Applications connect structured records, embeddings and full-text indexes through shared identifiers, and declare graph relationships in their schemas. Row, vector and full-text ingestion use separate APIs; the application supplies embeddings and coordinates these writes. Dedicated instances offer single-tenant compute in a chosen region, with optional asynchronous standby replication. Replication and recovery limits are detailed below and in the deployment documentation. ## Reading this website - [Full website text](https://originchaindb.com/llms-full.txt): generated from the current public pages during each build, with canonical source URLs. Includes product pages, documentation, examples, solutions, industries, articles and company information. - [Documentation home](https://originchaindb.com/docs): guides, data models, queries, integrations, operations and API reference. - [Current capabilities and answers](https://originchaindb.com/faq): concise answers linked to the relevant reference pages. - [Public sitemap](https://originchaindb.com/sitemap-index.xml): canonical public pages. Use the current API documentation for supported behavior and limitations, and the pricing page for current availability. Blog posts and release notes describe their publication dates; historical prices or capabilities are not current offers. Benchmark numbers apply to the documented workload and measurement conditions. Console screenshots marked as previews contain sample data. Cite the individual source URL when using information from the full-text file. ## Explore the product - [Product overview](https://originchaindb.com/product): four data models and their shared foundation. - [Database](https://originchaindb.com/product/database): how records, vectors, relationships and text indexes fit together. - [Authentication](https://originchaindb.com/product/authentication): identities, credentials and access controls. - [Storage](https://originchaindb.com/product/storage): persistence, storage configuration and recovery boundaries. - [SQL](https://originchaindb.com/product/sql): records, joins, aggregates, and query plans. - [Vector](https://originchaindb.com/product/vector): embedding similarity and metadata filters. - [Graph](https://originchaindb.com/product/graph): connected records, neighborhoods, and paths. - [Full-text](https://originchaindb.com/product/fts): ranked keyword and phrase retrieval. - [Ask](https://originchaindb.com/product/ask): natural-language query compilation and plan inspection. - [Solutions](https://originchaindb.com/solutions): search and RAG, agent memory, recommendations, and other application patterns. - [Resources](https://originchaindb.com/resources): documentation, benchmarks, and engineering insights. ## Docs - [Quickstart — create a database and run a query](https://originchaindb.com/docs/quickstart): connection details, a schema, records and a first query. - [HTTP API reference](https://originchaindb.com/docs/api): All `/v1/*` endpoints — schemas, rows, query, sql, vector, fts, graph, migrations, replication. - [Define schemas (TOML manifests)](https://originchaindb.com/docs/schemas): primary_key, columns, indexes, relations. - [Insert data](https://originchaindb.com/docs/insert): rows, batches and data ingestion. - [Query data](https://originchaindb.com/docs/query): plan tree, SQL, ask. - [SQL reference](https://originchaindb.com/docs/sql): JOINs, aggregates, CTEs, planner traits. - [Oracle SQL dialect](https://originchaindb.com/docs/compatibility/oracle-dialect): sequences, NVL, FROM DUAL, MINUS, CONNECT BY, ROWNUM; reachable through a PostgreSQL driver. No TNS listener - an Oracle client library cannot connect. - [Vector search reference](https://originchaindb.com/docs/vector): HNSW, fast vs. high_recall modes, recall@10 measurements. - [Full-text search reference](https://originchaindb.com/docs/fts): BM25, phrase search and explicit text indexing. - [Graph reference](https://originchaindb.com/docs/graph): BFS, Dijkstra, reverse traversal. - [Natural-language `/ask`](https://originchaindb.com/docs/ask): rule compiler + LLM fallback, plan-only mode. - [SDKs](https://originchaindb.com/docs/sdk): TypeScript, Python and Go. - [SQL client connection](https://originchaindb.com/docs/connect-sql-client): PostgreSQL-compatible connections and the supported SQL surface. - [Elasticsearch-compatible API](https://originchaindb.com/docs/connect-elasticsearch): connecting an Elasticsearch client to supported bulk, Query DSL and aggregation operations; not the full Elasticsearch product. - [Elasticsearch API reference](https://originchaindb.com/docs/elasticsearch): indexing, search, aggregation and compatibility limits. - [Connect over the MySQL and SQL Server wire protocols](https://originchaindb.com/docs/connect-wire-protocols): connection shape, TLS requirements, shared and per-user credential models, verified client examples, and the measured limits of the MySQL-wire and TDS listeners. - [Deploy](https://originchaindb.com/docs/deploy): dedicated configurations - single-tenant compute, region isolation. - [MySQL wire protocol compatibility](https://originchaindb.com/docs/compatibility/mysql-wire): how an instance speaks the MySQL wire protocol - connection shape, TLS, credential models, the catalog surface, and the measured limits. Existing MySQL clients connect; OriginChainDB does not distribute MySQL software. - [Operate (ops)](https://originchaindb.com/docs/ops): backups, PITR, replication. - [SQL Server wire protocol (TDS) compatibility](https://originchaindb.com/docs/compatibility/sqlserver-tds): SQL Server drivers connect, authenticate and query over TDS. The T-SQL procedural language is not implemented - the translated set, the refused set, and what a migration really costs. ## Examples (copy-paste JSON) - [All examples](https://originchaindb.com/docs/examples): index of every example page. - [SQL examples](https://originchaindb.com/docs/examples/sql) - [MySQL wire examples](https://originchaindb.com/docs/examples/mysql) - [Vector examples](https://originchaindb.com/docs/examples/vector) - [SQL Server wire (TDS) examples](https://originchaindb.com/docs/examples/sqlserver) - [Full-text examples](https://originchaindb.com/docs/examples/fts) - [Elasticsearch API examples](https://originchaindb.com/docs/examples/elasticsearch) - [Graph examples](https://originchaindb.com/docs/examples/graph) - [Natural-language examples](https://originchaindb.com/docs/examples/ask) - [Multi-model ingestion example](https://originchaindb.com/docs/examples/atomic-multi-shape): consult the endpoint contracts; ordinary row writes do not automatically supply embeddings or full-text documents. - [Errors](https://originchaindb.com/docs/examples/errors) ## Architecture & engineering posts - [Architecture overview](https://originchaindb.com/architecture): four data models, separate ingestion paths, application-coordinated retrieval, storage and dedicated writer/standby topology. - [Industries](https://originchaindb.com/industries): banking, payments, lending, capital markets, wealth management, mutual funds, insurance, NBFCs, aviation, healthcare, e-commerce, contact centres and telecom. - [Pricing](https://originchaindb.com/pricing): free database, pooled configurations and dedicated estimates. Configuration · 25 GB pooled is $29/month with 20 GB monthly data transfer; Configuration · 50 GB pooled is $49/month with 50 GB monthly data transfer. Both can be selected in the console; review checkout before paying. Monthly billing in USD. These are database configurations, not the SQL Pro add-on. - [Engineering blog (RSS)](https://originchaindb.com/api/blogs/rss.xml): full feed. - [Engineering journal](https://originchaindb.com/blogs): dated articles with diagrams, examples and original authorship. - [Benchmarks](https://originchaindb.com/benchmarks): published datasets, results, commands and measurement scope. ## Application patterns - [Search and RAG](https://originchaindb.com/solutions/rag): retrieve passages and retain their source context. - [AI agents](https://originchaindb.com/solutions/agents): application state, memory retrieval and connected records. - [Recommendations](https://originchaindb.com/solutions/recommendations): candidates, relationships and application ranking. - [Personalization](https://originchaindb.com/solutions/personalization): preferences, interactions and content retrieval. - [Fraud detection](https://originchaindb.com/solutions/fraud-detection): connected evidence for application-defined investigations. - [Real-time analytics](https://originchaindb.com/solutions/real-time-analytics): query operational records and inspect their context. ## Key facts for retrieval - **Data models:** SQL records, vectors, full-text indexes and graph relationships use one database engine. Ask is a query interface, not a fifth storage model. - **Ingestion:** records, vectors and full-text documents use separate requests. Applications generate embeddings, coordinate writes and combine retrieval results. Shared IDs connect indexes back to source records. - **Vector search:** available index types and search settings are documented in the vector reference. Compare recall and latency using the published workload and methodology; in-process results are not managed HTTP latency guarantees. - **Write durability:** consult the operation and deployment documentation for acknowledgement, secondary-index persistence and recovery behavior. Do not infer a cross-model snapshot or a multi-request transaction from a successful row write. - **Replication:** dedicated primary/standby configurations replicate asynchronously. Acknowledgement does not wait for the standby. Promotion is configuration-dependent; automatic promotion is opt-in and off by default. Consult the operations guide for lease, lag and recovery conditions. This is not synchronous quorum commit or a zero-data-loss guarantee. - **Backups:** dedicated recovery options and current preview limitations are documented in the operations guide. Snapshot recovery, change archiving and restore-to-timestamp have different prerequisites. Free and pooled configurations do not include point-in-time recovery. - **Deployment model:** managed cloud. Free uses shared, scale-to-zero infrastructure. Upcoming pooled configurations use shared compute capped at 1 vCPU, a single node in India (Mumbai), and no adjustable hot-data memory, high availability, replicas or additional nodes. Dedicated configurations run on virtual machines that serve only that one customer. In-place changes between free, pooled and dedicated are not available yet. - **Legal entity:** Silicoyn Technologies Pvt Ltd. ## Optional - [Changelog](https://originchaindb.com/changelog) - [Company](https://originchaindb.com/company): the team and legal entity behind OriginChainDB. - [Events](https://originchaindb.com/events): conferences, meetups and webinars we speak at, exhibit at or host. - [Website status](https://originchaindb.com/status): website build metadata and browser checks, not database instance uptime. - [Security](https://originchaindb.com/security) - [Privacy](https://originchaindb.com/privacy) - [Terms](https://originchaindb.com/terms) - [Cancellation & Refund Policy](https://originchaindb.com/refund-policy) - [DPA](https://originchaindb.com/dpa) - [Contact](https://originchaindb.com/contact) - [FAQ](https://originchaindb.com/faq) ## News and careers - [Newsroom](https://originchaindb.com/newsroom): OriginChainDB announcements, product updates, collaborations and industry perspectives. External source headlines are labelled separately from company news. - [Newsroom RSS](https://originchaindb.com/newsroom/rss.xml): OriginChainDB announcements and original industry perspectives. - [Careers](https://originchaindb.com/careers): no open roles currently; expressions of interest welcome. # Full public content Extracted from the canonical rendered pages listed below. Headings, code examples, tables and source links are retained. Examples are illustrative and may require setup. Historical articles, benchmarks and release notes retain their original scope; consult current documentation for supported behavior. This file does not report live service status. Coverage and content hashes: https://originchaindb.com/llms-manifest.json --- # OriginChainDB — The distributed multimodal database Canonical source: https://originchaindb.com/ Sitemap last modified: 2026-09-28T07:29:43.000Z Distributed multimodal database # One database. Four data models. Build search, agents, and connected applications with SQL, vector similarity, graph traversal, and full-text search in one managed engine. [Start free](https://app.originchaindb.com/signup) [See how it works](https://originchaindb.com/#how-it-works) Start freeScale with dedicated resources SQLStructured records VectorSemantic similarity GraphConnected context Full-textPrecise retrieval One engine four data models Why OriginChainDB ## Your application needs more than one kind of query. One engine to work with Keep structured records, vectors, relationships, and search indexes in one database. Context you can connect Use record identifiers to connect a search result to its source and related data. Your application stays in control Choose the queries, combine the results, and apply the rules that matter to your product. 01 / Four data models ## Give each question the right query. Start with the model you need. Add other query types as your application grows. [SQL ### Query your records. Filter rows, join tables, and aggregate the data your application runs on. 01 / Explore SQL](https://originchaindb.com/product/sql)[Vector ### Find similar meaning. Search your embeddings for similar items, passages, and memories. 02 / Explore Vector](https://originchaindb.com/product/vector)[Graph ### Follow relationships. Traverse relationships to reveal the context between your records. 03 / Explore Graph](https://originchaindb.com/product/graph)[Full-text ### Match the words that matter. Retrieve relevant documents with ranked full-text and phrase search. 04 / Explore Full-text](https://originchaindb.com/product/fts) Or ask in your own words. Natural-language queries with `/ask`.[Explore natural-language queries](https://originchaindb.com/product/ask) 02 / How it works ## See how the pieces connect. One support manual. Four useful ways to find the answer. Knowledge searchInteractive illustration 01 Source record `passage:07` ### AX-7 service manual Replace the battery when capacity drops. Disconnect power before servicing. Product AX-7 Version 3 02 Choose a query 03 Inspect the evidence SQL · matching record Which passages cover AX-7? `passage:07` ### Battery replacement Product AX-7 · version 3 Read the source row behind the result. Prepared with separate writes. Your app writes records and text/vector indexes, supplies embeddings, and combines retrieval results. Illustrative records. Choose a model to explore its role. [Build this workflow](https://originchaindb.com/docs/rag) 03 / Inside the console ## All your data. One place to work. Inspect rows, explore embeddings, trace connections, and read search results in the same workspace. [SQL](https://originchaindb.com/product/sql)[Vector](https://originchaindb.com/product/vector)[Graph](https://originchaindb.com/product/graph)[Full-text](https://originchaindb.com/product/fts) OriginChainDB console Write SQL. Inspect results. Understand the plan. [Explore SQL](https://originchaindb.com/product/sql) [Image: OriginChainDB SQL workbench with a SELECT query and five sample document rows.] Actual console · Sample data · Swipe to explore [View full size ↗](https://originchaindb.com/_astro/sql-dark.B_81ElDM.png) 04 / Built for developers ## Start with a query. Go from a table to an application result. Use SQL directly or send it over the HTTP API. ``` SELECT id, title FROM public.documents LIMIT 3; ``` Example result3 illustrative rows Illustrative records from a populated public.documents table; no query is executed on this page. | id | title | | --- | --- | | `doc-001` | Getting started | | `doc-002` | Vectors and similarity | | `doc-003` | Relationships in your data | Example only. No request is sent. 1. ### Get your connection details. Set `OC_URL`, `OC_TENANT` and `OC_TOKEN` from your instance. [Open the quickstart](https://originchaindb.com/docs/quickstart) 2. ### Create and populate the table. This example requires `public.documents` with `id` and `title` columns. [Create your first records](https://originchaindb.com/docs/sql/quickstart) 3. ### Read your data. Run the SQL in your workbench or send the HTTP request from your application. [Explore the query API](https://originchaindb.com/docs/api#sql) First-party SDKs Connect your tools PostgreSQL wire: limited preview. Elasticsearch: supported API subset. 05 / Deployment & operations ## Start with an idea. Build for production. A managed database, in your region. Start free, then choose dedicated resources for your workload. Your region. Your resources.Dedicated configurations run on single-tenant compute. Connected across nodes.Add a warm standby with asynchronous replication. Plan access and recovery.Configure access controls and daily snapshots for dedicated instances. [Explore deployment options](https://originchaindb.com/docs/deploy) YOUR APPLICATION Your chosen region OriginChainDBPrimary engine SQLVectorGraphFull-text Asynchronous replication Warm standbyOptional Illustrative dedicated topology [Read the details ↗](https://originchaindb.com/docs/deploy#replication) 06 / Built around your workload ## What will you build with it? [Explore all solutions](https://originchaindb.com/solutions) [RETRIEVE ### Search & RAG Retrieve passages by meaning and keywords, then give your model source-backed context. Vector / Full-text / SQL](https://originchaindb.com/solutions/rag)[REMEMBER ### AI agents Keep run state, retrieve useful memories, and trace the records behind each step. SQL / Vector / Graph](https://originchaindb.com/solutions/agents)[DISCOVER ### Recommendations Find similar items and connected interests. Apply your own eligibility and ranking rules. Vector / Graph / SQL](https://originchaindb.com/solutions/recommendations)[INVESTIGATE ### Fraud investigation Follow shared accounts, devices, and transactions to inspect connected activity. Graph / SQL](https://originchaindb.com/solutions/fraud-detection)[ADAPT ### Personalization Use session preferences and relevant content to shape what your application shows next. SQL / Vector / Full-text](https://originchaindb.com/solutions/personalization)[UNDERSTAND ### Operational analytics Group orders, count events, and inspect the records behind your application's metrics. SQL](https://originchaindb.com/solutions/real-time-analytics) 07 / Make an informed choice ## Look inside. Then put it to work. Review the architecture, test your workload, and know what your application will own. [01Architecture ### Understand what runs underneath. Read about the engine, indexes, persistence, and replication.](https://originchaindb.com/architecture)[02Benchmarks ### Inspect the test before the number. Check workloads, configurations, and methodology before comparing results.](https://originchaindb.com/benchmarks)[03Compatibility ### Check the queries your app needs. Review supported SQL and adapter behavior before planning an integration.](https://originchaindb.com/docs/connect-wire-protocols) ### A few things worth knowing. Have a specific workload? [Talk it through with us](https://originchaindb.com/contact) What does multimodal mean here? Four ways to work with data in one engine: SQL for records, vector search for similarity, graph traversal for relationships, and full-text search for words. Natural-language /ask is another query interface. [Explore the database](https://originchaindb.com/product/database) Does adding a record also create its embedding and search index? Your application generates embeddings and uses the dedicated indexing APIs. Shared identifiers connect the results to their source records. Your application coordinates these writes and combines query results. [Read the vector data model](https://originchaindb.com/docs/schemas/vector) Can I use my existing database tools? Use the HTTP API and SDKs, or supported wire and Elasticsearch adapters. Compatibility depends on the enabled adapter and its supported operations; check your queries against the integration guides. [Find an integration](https://originchaindb.com/docs/sdk) How do I move from an experiment to production? Start on the free tier, then choose a dedicated configuration for your workload and region. Review access controls, capacity, snapshots, and the optional asynchronous standby before launch. [Plan your deployment](https://originchaindb.com/docs/deploy) From the Newsroom ## Inside OriginChainDB. Product news, collaborations and our perspective on what comes next. [Explore the Newsroom ↗](https://originchaindb.com/newsroom) From your first query to your next application. ## Bring your data. Build your next idea. [Start free](https://app.originchaindb.com/signup)[Talk to our team](https://originchaindb.com/contact) [Or run your first query ↗](https://originchaindb.com/docs/quickstart) --- # How OriginChainDB works | OriginChainDB Canonical source: https://originchaindb.com/architecture Sitemap last modified: 2026-09-23T08:33:23.000Z Architecture / Distributed multimodal database # Four data models. See how they connect. SQL, vector, graph and full-text in one engine. Follow a record from your application to its indexes, query results and deployment. [Build your first workflow](https://originchaindb.com/docs/quickstart)[Explore the database](https://originchaindb.com/product/database) The design in three lines 1. 01 One engineRecords, embeddings, relations and searchable text. 2. 02 Explicit interfacesChoose how to write and query each representation. 3. 03 Your deploymentCompute, replication and recovery for your workload. Inside the system ## Follow a write. Trace a read. Switch views to see where each responsibility lives. Follow the data Illustrative architecture Your application ### An article, ready to index. Prepare fields, an embedding and searchable text. `article:07` Separate authenticated requests OriginChainDB engineSame tenant · shared record ID Row write `id: article:07title: Battery guideauthor: person:03` SQL + Graph Fields and declared relations Vector put `id: article:07vector: [0.2, 0.8, …]`You supply the embedding Vector index Embedding and metadata Full-text index `id: article:07text: Replace the battery…`Choose fields to search Full-text index Terms and document references A row write does not also submit its embedding or text. Your application coordinates these calls and their retries. Three views of the system, with the application and engine responsibilities shown separately. The query layer ## Use the model that fits the question. A shared record ID connects the representations. Each query interface serves a different kind of question. 01 / SQL ### Find and shape records. Filter fields, join related tables and aggregate results. Start with a registered schema and declared indexes. `Which articles are published?`[SQL quickstart](https://originchaindb.com/docs/sql/quickstart) 02 / Vector ### Search by similarity. Store embeddings from your chosen model. Query nearest neighbors with the same dimensions and apply metadata filters. `Which articles are similar?`[Vector quickstart](https://originchaindb.com/docs/vector/quickstart) 03 / Graph ### Follow connections. Declare relations and traverse neighbors or paths. Use relationships to explore the context around a record. `Who wrote this article?`[Graph quickstart](https://originchaindb.com/docs/graph/quickstart) 04 / Full-text ### Find the right words. Index selected text fields and query words or phrases. BM25 ranking provides scored document matches. `Where is “battery replacement”?`[Full-text walkthrough](https://originchaindb.com/docs/fts#walkthrough) The write path ## One ID. Separate, deliberate writes. For the native APIs, saving a row does not also create its embedding or full-text entry. [See the three-call example](https://originchaindb.com/docs/examples/atomic-multi-shape/kb-article) 1. 1 ### Define the record Register fields, a primary key and relations. Use that record ID in each index entry. 2. 2 ### Prepare the other representations Your application generates embeddings and chooses searchable text. Submit these through their own endpoints. 3. 3 ### Keep updates coordinated Track which calls succeeded. Retry failed requests and update or remove related entries when source data changes. Retrieval ## Combine evidence. Keep the source. For a search or RAG workflow, query vector and full-text indexes separately, merge rankings in your application, then fetch source records by ID. Vector hybrid search combines dense and sparse vectors. Combining vector results with BM25 is a separate application workflow. [Build a retrieval workflow](https://originchaindb.com/docs/rag) An optional query interface ### Ask a question. Inspect the plan. `/ask` turns a supported natural-language request into an executable plan. It tries deterministic rules first, with a model fallback when needed. It is a query interface over the engine, not an extra storage model or a complete answer-generation pipeline. [Try natural-language queries](https://originchaindb.com/docs/nql/quickstart) Deployment ## Choose how your database runs. A dedicated deployment can pair a writer with a standby in the same region. Writer + optional standby ### Replication is asynchronous. The writer streams committed changes to a standby. Recent writes may be missing after an abrupt writer loss; promotion is operator-driven by default. The standby follows the writer; backups provide a separate path to restore data. [Explore deployment options](https://originchaindb.com/docs/deploy#replication) [Scaling across nodesQuery support across nodes See how the topology affects queries and transactions.](https://originchaindb.com/docs/multi-node#multi-node-limits)[Multiple regionsActive-active development Follow the published status and scope of cross-region writes.](https://originchaindb.com/docs/replication/multi-region) Operations ## Plan for the next write. And the next recovery. Monitor the writer, observe replication and practice restore. Each gives you a different view of database health. 01 ### Understand acknowledgments Learn what a successful write means on the primary, with endpoint-specific behavior in the durability guide. [Write durability](https://originchaindb.com/docs/ops#durability) 02 ### Restore from backups Restore from snapshots and available archives. The guide covers recovery points and current preview limitations. [Backup and restore](https://originchaindb.com/docs/ops#backups) 03 ### Observe the system Follow health, usage, write errors and replication lag. Use the operations guide to connect monitoring and incident response. [Operations guide](https://originchaindb.com/docs/ops#observability) Start building ## Start with a record. Add the queries you need. [Open the quickstart](https://originchaindb.com/docs/quickstart)[Read the API reference](https://originchaindb.com/docs/api) --- # Database benchmark results | OriginChainDB Canonical source: https://originchaindb.com/benchmarks Sitemap last modified: 2026-09-23T08:33:23.000Z [Resources](https://originchaindb.com/resources) Published benchmark report # Measurements. With the setup attached. Operational throughput, vector recall, and full-text relevance—read each result with its workload, hardware, and measurement boundary. Read the boundary first ## These runs measure different paths. An HTTPS result includes network time. An in-process result does not. Compare workloads within their stated conditions. What each measurement includesPublished test setups 01 / Remote HTTPSYCSB + BEIR 1. Client laptopIndia 2. HTTPS round trip~100 ms RTT to Mumbai 3. Reference engine2 vCPU · 2 GB 02 / In-region HTTPSYCSB sweep 1. Regional driverSame region as engine 2. HTTPS requestClient concurrency recorded 3. Configured tenant4 or 8 vCPU 03 / In-processANN-Benchmarks 1. Rust testDevelopment machine 2. Index searchSingle thread · AVX2 3. Recall + timingsRelease build · no HTTP The reference engine used substrate compression and a bounded memory guard. The in-region runs use separate production configurations. These published results do not establish an unthrottled engine ceiling or a service-level guarantee. 01 / Operational workloads ## YCSB: reads and writes over HTTPS. Five reported workloads: A, B, C, D, and F. Zipfian access on a 20K-row table, with 50,000 operations across the workloads. MeasurementRemote HTTPSLatency in milliseconds YCSB · reference engine, remote client | Workload | Mix | ops / sec | p50 (ms) | p95 (ms) | p99 (ms) | errors | | --- | --- | --- | --- | --- | --- | --- | | A | 50% read / 50% update | 292 | 104 | 143 | 202 | 0 | | B | 95% read / 5% update | 321 | 92 | 120 | 158 | 0 | | C | 100% read | 146 | 125 | 557 | 1220 | 0 | | D | 95% read latest / 5% insert | 338 | 92 | 119 | 138 | 0 | | F | read-modify-write | 215 | 165 | 225 | 252 | 0 | All five reported workloads recorded zero errors and no rate-limit hits. Workload E is not included in this dataset. ### In-region sweep Measured 11 June 2026 The same YCSB driver ran from a machine in the engine’s region against production-configuration tenants. The notes below are retained from the published run. YCSB · in-region client and provisioned tenant limits | Setup | ops / sec | p99 (ms) | Note | | --- | --- | --- | --- | | 4-vCPU node | 1,046 | - | at its provisioned ~1,000 req/s cap | | 8-vCPU node | 1,384 | 117 | single client; the driver is the limit, not the engine | | 8-vCPU node, 4-client fan-out (read) | 4,940 | 173 | 0 errors; the wall at ~5,000 is the provisioned cap, not an engine ceiling | The four-client row reports read traffic. A dash means p99 was not supplied for that row. Provisioned rate limits and client concurrency are part of this setup. Reproduction command from the report ``` python benchmarks/ycsb/run.py --record-count 20000 --op-count 10000 --workloads C,B,A,D,F ``` 02 / Vector retrieval ## Recall and query rate, together. SIFT 128-dimensional vectors, a 100K-vector subset, and 1,000 queries. HNSW with M=16 and ef_construction=200, using a single-threaded AVX2 kernel. MeasurementIn-processLatency in microseconds ANN-Benchmarks protocol · SIFT 100K subset | ef_search | recall@10 | QPS | p50 (µs) | p99 (µs) | | --- | --- | --- | --- | --- | | 10 | 0.286 | 187.9 | 5,208 | 8,170 | | 20 | 0.115 | 154.8 | 6,279 | 9,685 | | 50 | 0.232 | 110.7 | 8,966 | 11,313 | | 100 | 0.378 | 73.9 | 13,183 | 19,852 | | 200 | 0.565 | 47.6 | 20,718 | 26,153 | | 400 | 0.765 | 28.4 | 34,952 | 43,987 | | 800 | 0.918 | 16.2 | 61,570 | 76,478 | The published rows are shown without smoothing or interpolation, including the non-monotonic recall at ef_search 10 and 20. Choose a recall target, then compare the query rate and tail latency at that operating point. Reproduction scope ### Record the exact machine and test. This report identifies a Rust integration test on a development machine but does not include its invocation. The SIFT-1M path uses `SIFT_DATA_DIR` and was described for a 32 GB machine; no SIFT-1M result is published here. Large-corpus scope ### No measured 100M result is published. IVF and IVF-PQ are separate index choices. This HNSW dataset does not measure their memory use, latency, or recall at larger corpus sizes. [Review index choices](https://originchaindb.com/docs/vector) 03 / Full-text relevance ## BEIR: SciFact retrieval quality. NDCG@10 measures graded relevance in the first ten results. The Lucene BM25 reference is the baseline reported in the BEIR paper, Table 2. Published corpusSciFactOne dataset from the BEIR suite SciFact · OriginChainDB and published Lucene BM25 baseline | Dataset | Docs | Queries | OC NDCG@10 | Lucene BM25 | Δ | | --- | --- | --- | --- | --- | --- | | SciFact | 5,183 | 300 | 0.662 | 0.665 | -0.003 | The report used BM25 with default Anserini parameters and English tokenisation. The reported score difference is −0.003; the table does not compare every BEIR dataset. Indexing rate 73 docs/sec Query rate 67 q/s Reproduction command from the report ``` python benchmarks/beir/run.py --dataset scifact ``` Make a comparison you can use ## Repeat the workload that matters to you. Run the listed commands in a checkout containing the benchmark harness. Keep the setup and the raw output alongside every result. 01 ### Match the conditions. Record corpus size, query mix, concurrency, hardware, engine version, and network location. Keep provisioned limits visible. 02 ### Compare quality with speed. For vector retrieval, report recall with latency and QPS. For full-text, retain relevance judgments and the dataset split. 03 ### Retain the full result. Include tail latency, error counts, configuration, and run logs. Do not extrapolate a result to unmeasured workloads or corpus sizes. These are published run results, not current service telemetry. The tables cover only the configurations and datasets shown above. [Plan a workload evaluation](https://originchaindb.com/pilot) Bring your workload ## Put the database through your own test. [Follow the database quickstart](https://originchaindb.com/docs/quickstart)[Plan an evaluation](https://originchaindb.com/pilot) --- # OriginChainDB Blog — engineering, guides & ideas Canonical source: https://originchaindb.com/blogs Sitemap last modified: 2026-09-23T08:33:23.000Z The OriginChainDB journal [Follow via RSS](https://originchaindb.com/api/blogs/rss.xml) # Inside the database. Beyond the query. How we build it, what we learn, and the decisions behind your data. [Browse all articles](https://originchaindb.com/blogs#articles) Featured readJul 14, 2026 ## [How a graph database works vs joins](https://originchaindb.com/blogs/how-graph-databases-work) Follow the relationships behind connected data, and see where graph traversal and SQL joins fit. OriginChainDB[Read article](https://originchaindb.com/blogs/how-graph-databases-work) Keep exploring [Jul 14, 2026 · EngineeringHow vector search works: HNSW explained](https://originchaindb.com/blogs/how-vector-search-works)[Jul 8, 2026 · ArchitectureThe token economy and your AI bill](https://originchaindb.com/blogs/token-economy-occi) The reading room ## Find your next idea. 29 articles EngineeringJul 14, 2026 ### [How a graph database works vs joins](https://originchaindb.com/blogs/how-graph-databases-work) graph / cypher EngineeringJul 14, 2026 ### [How vector search works: HNSW explained](https://originchaindb.com/blogs/how-vector-search-works) vector-search / embeddings ArchitectureJul 8, 2026 ### [The token economy and your AI bill](https://originchaindb.com/blogs/token-economy-occi) ai-costs / token-economy ArchitectureJul 1, 2026 ### [Why the database is now multi-modal](https://originchaindb.com/blogs/multi-modal-is-the-new-foundation) multi-modal / ai EngineeringJun 22, 2026 ### [One database, every query shape](https://originchaindb.com/blogs/one-database-every-query-shape) database / ai EngineeringJun 7, 2026 ### [Multi-region cold standby at $5/month](https://originchaindb.com/blogs/multi-region-cold-standby-5-vs-78-per-month) dr / infrastructure EngineeringJun 6, 2026 ### [Fuzzing a database: 850,000 probes/day](https://originchaindb.com/blogs/850000-random-api-probes-a-day) fuzzing / reliability EngineeringJun 6, 2026 ### [Window functions, correlated subqueries](https://originchaindb.com/blogs/window-functions-correlated-subqueries-590-tests) sql / window-functions EngineeringMay 26, 2026 ### [RAG at 10M documents: version skew](https://originchaindb.com/blogs/rag-at-10m-documents-the-four-vendor-failure-mode) rag / vector-search PerformanceMay 12, 2026 ### [From a 20-second dashboard to 300 ms](https://originchaindb.com/blogs/from-20s-dashboard-to-300ms) infrastructure / performance TutorialsMay 12, 2026 ### [Automatic Idempotency-Key in the SDK](https://originchaindb.com/blogs/auto-idempotency-key-sdk) sdk / dx TutorialsMay 6, 2026 ### [Agent memory economics: 100M tool calls](https://originchaindb.com/blogs/agent-memory-economics) cost / agent-memory ArchitectureMay 6, 2026 ### [Write backpressure: 429 + Retry-After](https://originchaindb.com/blogs/backpressure-429) backpressure / rate-limiting ComparisonsMay 6, 2026 ### [OriginChainDB vs Redis: when each fits](https://originchaindb.com/blogs/vs-redis) redis / comparison ComparisonsMay 6, 2026 ### [OriginChainDB vs Supabase: when each fits](https://originchaindb.com/blogs/vs-supabase) supabase / comparison TutorialsMay 5, 2026 ### [AI feature cost at 100K to 10M users](https://originchaindb.com/blogs/ai-feature-cost-model) pricing / cost TutorialsMay 5, 2026 ### [Idempotent tool calls for agent loops](https://originchaindb.com/blogs/idempotent-tool-calls) agents / tool-calls ArchitectureMay 5, 2026 ### [Do you need a separate vector database?](https://originchaindb.com/blogs/no-separate-vector-database) architecture / vector-database TutorialsMay 5, 2026 ### [Per-key TTL: agent memory that forgets](https://originchaindb.com/blogs/per-key-ttl-agent-memory) ttl / agent-memory TutorialsMay 5, 2026 ### [RAG latency budget: where the time goes](https://originchaindb.com/blogs/rag-latency-budget) rag / performance ComparisonsMay 5, 2026 ### [OriginChainDB vs DynamoDB: when each fits](https://originchaindb.com/blogs/vs-dynamodb) dynamodb / comparison ComparisonsMay 5, 2026 ### [OriginChainDB vs Pinecone: when each fits](https://originchaindb.com/blogs/vs-pinecone) pinecone / vector-database ComparisonsMay 5, 2026 ### [OriginChainDB vs Postgres + pgvector](https://originchaindb.com/blogs/vs-postgres-pgvector) postgres / pgvector ProductMay 4, 2026 ### [Our depth-first roadmap to 1.0](https://originchaindb.com/blogs/depth-first-roadmap-to-1-0) roadmap / product PerformanceMay 4, 2026 ### [HA snapshot bootstrap: the cutover gap](https://originchaindb.com/blogs/ha-snapshot-bootstrap) ha / replication TutorialsMay 4, 2026 ### [OriginChainDB quickstart in five minutes](https://originchaindb.com/blogs/quickstart) tutorial / quickstart EngineeringMay 4, 2026 ### [Eight comparison pages in one afternoon](https://originchaindb.com/blogs/eight-vs-pages-three-agents-one-afternoon) engineering / ai-tools TutorialsMay 4, 2026 ### [OpenAPI spec and the AI coding loop](https://originchaindb.com/blogs/openapi-spec-and-the-ai-coding-loop) openapi / ai-tools EngineeringMay 4, 2026 ### [One MCP server for every AI IDE](https://originchaindb.com/blogs/mcp-server-launch) mcp / ai-tools Keep the useful bits close ## From a good read to a working query. Explore the examples in our docs, or follow new writing in your feed reader. [Open the docs](https://originchaindb.com/docs)[Subscribe via RSS](https://originchaindb.com/api/blogs/rss.xml) --- # Fuzzing a database: 850,000 probes/day - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/850000-random-api-probes-a-day Published: 2026-06-06T13:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal fuzzingreliabilitycanaryengineeringqa # Fuzzing a database: 850,000 probes/day A 24/7 fuzz canary fires ~10,000 random API probes a minute at a live engine. The 14 generator families, the invariants it asserts, and what it misses. OriginChainDB engineering Jun 6, 2026 About 8 min read Testing / Fuzzing ## Turn unexpected input into a reproducible case. 1. Generate Vary payloads and sequences 2. Exercise Call the database API 3. Observe Check defined invariants 4. Reduce Keep a minimal failing case A fuzzing loop explores inputs, checks invariants and turns failures into cases that can be investigated and tested again. TL;DR — A dedicated host in our managed cloud fires ~10,000 random API probes at the engine every minute, 24/7. That’s ~850,000 probes/day, ~310M/year. Every probe is recorded; every failure pages on-call within two minutes. Here’s what’s actually in the loop, how we assert “no bug” without ground-truth, and the surface we don’t yet cover. You can’t unit-test your way to production confidence. Unit tests prove a function does what you wrote it to do. They don’t prove the HTTP layer above it stays well-formed under unexpected input, the auth wall holds against omitted headers, or the planner doesn’t 500 on a SQL string a customer types by accident. The thing that does prove those properties is a fuzzer that runs forever against a real engine. We built one. It runs as a continuous loop on a dedicated host in our managed infrastructure, hammering the engine over its public HTTPS endpoint with about 10,000 randomly-generated probes per iteration, sleeping 60 seconds, and starting again. The loop has been running uninterrupted, on a real tenant, with real auth, for weeks. ## The numbers - 10,000 probes per iteration. One iteration = one invocation of the fuzzer with a fresh seed. - ~60 seconds per iteration. Run time + sleep. p99 latency across the run lands around 115 ms. - ~1,440 iterations/day. 10,000 × 1,440 = ~14M probes a day at full throttle. We currently throttle the canary at one iteration per minute to keep p99 anchored, which gives the ~850,000-probes-a-day floor we quote on call. - 14 generator families in the rotation, weighted so the cheap fixed-shape probes (`/health`, `/capabilities`) anchor latency while the heavier surfaces (SQL, vector, FTS, graph, NL ask) get the bulk of the coverage budget. Surface coverage across the engine’s public endpoints sits around 85%. - One metric, one alarm, one paging topic. The metric is `OriginChain/Fuzz/FuzzFailures`. The alarm fires on `>0` for two consecutive minutes. The paging topic is the same one routing every other engine alarm to on-call. The loop is a small script, supervised so it restarts on failure with a start limit (5 restarts per 60 seconds) so a runaway can’t take the host down. The loop body is short: ```bash while true; do LOOP_COUNTER=$((LOOP_COUNTER + 1)) SEED=$(( $(date +%s%N) ^ LOOP_COUNTER )) fuzz run \ --base-url "$BASE_URL" \ --bearer "$BEARER" \ --tenant "$TENANT" \ --iterations 10000 \ --workers 16 \ --seed "$SEED" \ --report "$REPORT_PATH" FAILURES=$(jq -r '.summary.failed // 0' "$REPORT_PATH") publish_metric \ --metric-name "FuzzFailures" \ --value "$FAILURES" || true sleep 60 done ``` The metric publish is wrapped in `|| true` so a monitoring-side outage never stops the canary. The fuzzer’s own report file is the source of truth for what failed; the metric is just the cheap pointer for the alarm. ## The 14 generator families The fuzzer doesn’t randomly mash bytes. It generates probes by family — each family knows the shape of one engine endpoint and randomizes within that shape. The bag of families is round-robin’d with random jitter so a long run gets balanced coverage. The current rotation: - `Health` / `Capabilities` / `Version` / `Usage` — cheap fixed-shape probes. Anchor the p99 to the boring path so a regression in `/health` shows up in the same metric. - `AuthWall` — probes sent without `Authorization: Bearer`. Asserts 401, not 403, not 500. - `SchemaCrud` / `RowCrud` — schema and row CRUD against random table names, random column types, random literals, random typos. - `SqlSelect` — random SELECT shapes against real and made-up tables. Stress-tests the SQL translator. - `Vector` / `Fts` / `Graph` — the three other query shapes. Random k, random metric, random filter. - `Query` — the typed query API. Random combinations of where/order/limit. - `Migrations` — schema evolution probes. Random shape upgrades against random shape versions. - `Ask` — the natural-language endpoint. Random English-ish prompts that may or may not parse to a plan. A probe carries an HTTP verb, a path, an optional JSON body, the generator family that produced it, and an `intent` string (“malformed SQL”, “missing PK”) that surfaces on failure so triage doesn’t reverse-engineer the generator. The full probe is serialized into the report file so a replay command re-issues the exact probe against any environment, including a local engine for diagnosis. ## How we assert “no bug” without ground-truth This is the part most people skip when describing a fuzzer. You can’t write `assert_eq!(actual, expected)` against a random probe — you don’t know what the expected output is. What you do know is the invariants the endpoint must maintain regardless of input. The assertion rules, in order: 1. No 5xx. Ever. A 500 means the engine panicked or returned an unhandled error. Every 5xx is a bug. Generated probes that are deliberately malformed should return 400, not 500. This is the single highest-signal assertion in the loop. 2. The response body is parseable JSON. Even error responses. If the engine returns `Content-Type: application/json` and the body fails `serde_json::from_slice`, that’s a bug — it means a code path is emitting raw `format!` output where it should be emitting structured errors. 3. The auth wall holds. Probes with `bearer_omitted: true` must come back 401. Not 403, not 500, not 200. A single 200 on a no-auth probe is the kind of bug that ends careers; the canary fails immediately and pages. 4. The status code is within the expected band for the family. A `Health` probe returning 503 is OK (the engine is degraded). A `Health` probe returning 400 is a bug — health endpoints don’t take input. Each family carries its own band. Generated probes that are deliberately wrong are expected to fail with 4xx — that’s not a failure of the engine, it’s a successful auth/parse/validate gate. The assertion rules distinguish “the engine correctly rejected garbage” from “the engine accidentally accepted garbage” from “the engine crashed on garbage.” Only the last two count as failures. ## The p99 number, honestly The current loop reports p99 around 115 ms across the full probe mix. That number is dominated by the heavier endpoints — vector search and SQL — not the cheap `/health` probes. The fixed-shape probes anchor the median to the low single-digit ms range; the p99 is where the planner work, the index lookups, and the natural-language parse actually show up. We publish p99 and p99.9 into the per-iteration report. We watch both. The day p99 jumps to 200ms without a deployment, something has slipped in the planner — same dashboard you’d build to watch a real customer workload, except the canary’s “customer” is a 10,000-probe-a-minute random adversary. ## What we don’t cover Three surfaces the canary doesn’t touch, in order of how much we care: - Long-poll `/watch` subscriptions. The canary issues short-lived HTTP requests. The reactive `/watch` surface — the one that keeps a connection open and streams change frames — is fundamentally a different shape and isn’t in the rotation. We test it in unit tests and integration tests; it’s not in the live canary. - Destructive migration paths. The `Migrations` family covers shape-upgrade probes that the engine can recover from. It deliberately does not run `DROP SCHEMA` or unconstrained `DELETE`. The canary’s tenant is real; a destructive probe that fired in production would, by construction, destroy production data. The destructive paths are tested in a sandbox tenant in CI, not in the live loop. - State-aware behavior. Each probe is independent. The canary doesn’t write a row, read it back, assert equality. It asserts shape invariants on every response, not semantic equivalence across requests. Read-your-own-write consistency is unit-tested and integration-tested; the canary is the “no surprise crashes” layer above that, not the “the database returns correct answers” layer. These are real gaps. We don’t pretend they aren’t. ## The CI counterpart The same fuzz binary runs as a daily job in CI, against the same canary tenant, from outside our managed cloud. That’s the second opinion when the in-tenant loop’s metric drifts. If both the in-tenant loop and the external CI run agree, we trust the result. If they disagree, the disagreement is itself the signal — usually a routing or auth path that behaves differently from inside the cloud vs from outside. The CI workflow uses a credentials secret keyed off the same bearer the canary uses; rotating one rotates the other. ## Why the canary lives next to the engine It would be tempting to run the canary from outside the cloud entirely — that’s what the daily CI run does. We run the in-tenant loop next to the engine for one reason: latency. From inside the same region, the canary’s p99 number reflects the engine’s p99, not the engine plus public-internet network jitter. Putting the canary on the same network the engine serves from gives us a clean signal that a regression is in the planner or the index code, not in the network edge or load balancer. That’s load-bearing for the alarm threshold. A p99 jump on a remote canary could be a network anomaly. A p99 jump on a co-located canary, with the cheap fixed-shape probes still anchored to low-single-digit ms, is the engine. ## What the canary catches In the months it’s been running, the canary has caught: - A planner regression where a malformed `ORDER BY` returned 500 instead of 400. Found in iteration #3,481. Caught in two minutes. - An auth path where a specific bearer-token shape was 403’d by middleware before reaching the engine, instead of 401’d by the engine. The canary’s `AuthWall` family is what surfaced it. - A `/usage` endpoint regression where a specific tenant ID format returned a malformed JSON envelope. Caught by assertion rule #2. None of these were caught by unit tests. They were caught by the loop hammering shapes the test suite hadn’t thought to write. ## Try it on your own engine If you’re evaluating OriginChainDB and want to know what “stable enough to run an agent against” looks like, the canary is the operational answer. Create a [free database](https://originchaindb.com/docs/quickstart), point your own load at the engine, and watch the same metric the canary is publishing. The graph shape is the data. One endpoint. Random probes. Forever. No 5xx. --- # Agent memory economics: 100M tool calls - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/agent-memory-economics Published: 2026-05-06T10:26:12.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal costagent-memoryscalingttltutorial # Agent memory economics: 100M tool calls An agent at 100M tool calls/month writes ~50 GB of traces, but LLM inference is 99.8% of the bill. The four levers that actually move agent memory cost. OriginChainDB Team May 6, 2026 About 6 min read Agent systems / Cost ## Choose what your agent remembers. 1. Working context Keep the current task close 2. Useful history Retain selected tool results 3. Embeddings Track what needs re-embedding 4. Retention Expire data with a purpose Memory cost depends on what you keep, how long you keep it, and when you recompute representations. The article works through the assumptions. TL;DR - An autonomous AI agent with 100M tool calls per month produces ~50 GB of trace data and ~75 GB of embeddings. Storage is the small line item; what dominates is recompute cost (re-embedding stale state) and retention (keeping calls around long enough to be useful but not so long they become legal liability). ## The shape of the cost For an agent doing 100M tool calls per month: | Line item | Per call | Monthly total | Cost (entry configuration) | | --- | --- | --- | --- | | Tool-call trace (JSON) | ~500 bytes | 50 GB | ~$5 | | Embedding of trace | 768-dim f32 = 3 KB | 300 GB | ~$30 | | Re-embedding (stale window) | varies | varies | ~$50-200 (LLM API) | | Retention sweep | - | - | included | Total OriginChainDB footprint: ~$35/month at this scale. Total LLM/embedding spend: $50-200/month depending on the re-embedding cadence. For comparison, the LLM tool-call costs themselves (the actual inference for each call) at 100M calls × ~$0.0002 average per call = $20,000/month. So storage and embedding are 0.2% of the cost; the LLM is 99.8%. ## The lever that matters: retention The single largest cost lever isn’t storage rate - it’s how long you keep the data. A 100M-call/month agent over 12 months at full retention: 600 GB of traces + 3.6 TB of embeddings = ~$420/month in storage by month 12. The same agent with 90-day retention: 150 GB + 900 GB = ~$105/month, plateauing. The question is: do you actually need every tool call from 11 months ago? Usually not - most retrieval queries hit the recent window (last 7-30 days). After 90 days the value-per-byte drops sharply. [Per-key TTL](https://originchaindb.com/blogs/per-key-ttl-agent-memory) makes this trivial: tag every trace with a 90-day TTL on write; the substrate forgets it on its own. No cron job, no “DELETE WHERE created_at <…” that locks the table. ## The lever that doesn’t matter: storage rate People obsess over per-GB cost. For agent memory it’s a rounding error. Even at 10x our cost (some specialty databases charge $1-2/GB-month for managed storage), a 100M-call/month agent’s storage bill is ~$300/month. Compared to the $20K LLM spend, optimising the storage rate from $0.10 to $0.03/GB-month saves you ~$200/month - a 1% savings on the total bill. Spend the optimisation budget on the LLM line. ## The lever that’s invisible: re-embedding When you ship a new embedding model (or fine-tune the existing one), the old embeddings in storage no longer match query-time embeddings. You have two options: 1. Re-embed everything. 300M existing vectors × $0.00002 per embedding = $6,000 one-time. 2. Dual-version with gradual drift. Tag each embedding with model version; query-time embeddings use the same version; old vectors expire naturally via TTL; the new version takes over the corpus over your retention window. Option 2 is what we recommend. Combined with sensible retention, model upgrades become a no-op - you just start writing the new version and let TTL handle the old. The dual-version pattern needs the substrate to handle multi-version embeddings cleanly. OriginChainDB’s [shape versioning](https://originchaindb.com/docs/schemas/reference) does this - same shape, different version, queries scope themselves automatically. ## The lever that’s structural: what you store The biggest cost reduction comes from being honest about what you’re storing. Common bloat we see: - Storing the full LLM response, not just the tool-call decision. The reasoning trace is 10x the size of the tool decision and rarely retrievable enough to justify keeping. - Storing every intermediate state of the agent’s planning. Useful for debugging, expensive at scale. Move it to a debug-only sink with short TTL. - Storing duplicate embeddings. If the same input embeds to (approximately) the same vector, dedupe at write time via a hash key. The cleanest production agents store: - The tool decision (what was called, with what args, getting what result). - One embedding per intent, not per call. - A foreign key to the user/session/parent-trace. Everything else either lives in shorter-TTL debug storage or doesn’t get stored at all. ## A worked example Production agent doing 1M tool calls/day (30M/month): Naive storage: - Full reasoning + decision + result: 5 KB/call → 150 GB/month → $15 - Embedding of every call: 3 KB/call → 90 GB/month → $9 - Cumulative at 12 months retention: 1.8 TB + 1.1 TB = ~$290/month plateau Disciplined storage: - Decision only: 500 bytes/call → 15 GB/month → $1.50 - Embedding of distinct intents (~10% of calls): 0.3 KB/call effective → 9 GB/month → $0.90 - Cumulative at 90 days retention: 45 GB + 27 GB = ~$7/month plateau The same agent costs 40x less per month with disciplined storage. The LLM bill is unchanged either way. ## What changes at 1B calls/month Storage line scales linearly. LLM line scales linearly. Network egress (if you stream traces somewhere) becomes visible. Retention sweep becomes a real workload - at this volume, daily compaction matters; OriginChainDB handles it asynchronously and out of the foreground request path. The lever that becomes more important at this scale: embedding model choice. A 384-dim model uses half the storage of a 768-dim model and is often “good enough” for tool-call retrieval (which doesn’t need the same fidelity as document search). Switching halves the embedding storage line and roughly halves embedding-generation cost. ## FAQ ### What’s the per-call storage cost for agent memory? At our pricing, ~$0.0000005 per call for traces + ~$0.000001 per call for embeddings. Both rounding-error compared to the LLM cost. ### Should I re-embed when I switch models? Usually no - let TTL handle it. Dual-version embeddings during the transition window; old vectors expire naturally. One-shot re-embedding is only worth it if your retention window is very long and you can’t tolerate query-time inconsistency. ### How do I know when to expire traces? Look at your retrieval queries. If 90% of them only reach the last 30 days, set TTL to 90 days. If your debugging workflows need 6 months, set it to 6 months. Don’t pick a number from a blog post - measure your own access pattern. ### What’s the right embedding dimension for tool-call memory? Smaller than you think. 384 is fine for most cases. 768 is the default we recommend; 1536 (OpenAI’s text-embedding-3-large) is overkill for short tool calls. ### Can I get a cost breakdown for my own workload? We’re happy to model it - share your call rate, average payload size, and retention window and we’ll do the math. ## What to read next - [The cost of an AI feature at scale](https://originchaindb.com/blogs/ai-feature-cost-model) - the broader cost picture. - [Per-key TTL](https://originchaindb.com/blogs/per-key-ttl-agent-memory) - the retention primitive. - [Idempotent tool calls](https://originchaindb.com/blogs/idempotent-tool-calls) - pairs naturally with intent-deduplication. --- # AI feature cost at 100K to 10M users - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/ai-feature-cost-model Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal pricingcosttutorialscalingllm # AI feature cost at 100K to 10M users A per-user cost model for an AI feature: LLM tokens are 95% of the bill, embeddings 3%, the database 2%. Line items at 100K, 1M and 10M users. OriginChainDB Team May 5, 2026 About 6 min read Planning / AI economics ## Follow one request through the cost model. 1. Model tokens Input context and generated output 2. Embeddings New content and query vectors 3. Database Storage and query workload 4. Usage Active users × requests per user A cost model separates unit costs from request volume. This diagram identifies the inputs; it does not represent measured spend or current pricing. TL;DR - An AI feature’s bill is roughly 95% LLM tokens, 3% embedding generation, and 2% database. A personalized feed lands near $0.015 per user per month at 100K, 1M and 10M MAU alike, because the dominant line scales linearly with users. Pick the smallest model that works before you optimize anything else. ## Why this matters AI features look free in the prototype phase. You’re paying $20/month for an OpenAI key, $50/month for a small managed Postgres, and your “vector store” is a pickle file on disk. The unit economics work because there are no users. Then you ship. By the time you have 100K MAU, the math has shifted: the OpenAI bill is $5K, the database is $300, and the vector sidecar is another $200. By 1M MAU you’re at $50K/month if you didn’t think about caching. By 10M MAU it’s a real budget line. Most of that pain is avoidable with engineering decisions you make before you scale, not after. ## The model: a personalized content feature Each user gets a personalized “for you” feed: - 50 user actions per day generate signal (views, clicks, dismisses) - 1 personalization read per session, ~3 sessions per user per day - ANN search returns top 20 candidates from a 1M-item corpus - LLM re-ranks the top 20 to produce the final 10 - User profile embedding is refreshed once per day if signal accumulated That’s the per-user shape. Multiply by MAU. ## Per-user daily numbers | Operation | Count/day | Notes | | --- | --- | --- | | Signal writes | 50 | small JSON, 100 bytes | | Profile reads | 3 | 1KB per read | | Profile embeddings refresh | 1 | 768-dim f32 = 3KB | | ANN searches | 3 | top-20 from 1M corpus | | LLM re-rank calls | 3 | ~2K input tokens, ~200 output | | Hydrate top-10 reads | 30 | 30 batched gets | Storage growth: - 50 signal records/day × 100 bytes = 5KB/day per user - 1 embedding/day × 3KB = 3KB/day per user - ~8KB/day per user total ## Costs at 100K MAU | Line item | Volume | Unit cost | Monthly | | --- | --- | --- | --- | | LLM re-rank (Sonnet) | 9M calls | ~$0.0006 each | $5,400 | | LLM re-rank (Haiku) | 9M calls | ~$0.00015 each | $1,350 | | Embedding generation | 3M new vectors | ~$0.00002 each | $60 | | Database I/O (OriginChainDB, entry configuration) | All operations | flat monthly | ~$50 | | Database storage (24GB cumulative) | OriginChainDB, entry configuration | flat monthly | (included) | Totals: - With Sonnet: ~$5,500/month → $0.055 per user/month - With Haiku: ~$1,460/month → $0.015 per user/month Notice: the LLM is 95% of the bill. Storage is rounding error. Picking the smallest sufficient model is the highest-leverage decision you’ll make. ## Costs at 1M MAU Roughly 10x the volume. The interesting wrinkle is which line items scale linearly and which don’t. | Line item | Volume | Monthly | | --- | --- | --- | | LLM re-rank (Haiku) | 90M calls | ~$13,500 | | Embedding generation | 30M new vectors | ~$600 | | Database I/O (OriginChainDB) | Higher tier needed | ~$400 | | Database storage (240GB cumulative) | Higher tier | ~$200 | Total: ~$14,700/month → $0.015 per user/month The per-user cost doesn’t change much from 100K to 1M because the LLM dominates and scales linearly. This is good: you can model what 10M will cost without a calculator. The database lines do tier up - at 1M MAU you’re past the smallest plan. But the database is still <5% of the bill. ## Costs at 10M MAU | Line item | Volume | Monthly | | --- | --- | --- | | LLM re-rank (Haiku) | 900M calls | ~$135,000 | | Embedding generation | 300M new vectors | ~$6,000 | | Database I/O | Enterprise tier | ~$2,500 | | Database storage (2.4TB) | Enterprise tier | ~$1,500 | Total: ~$145,000/month → $0.0145 per user/month At this scale, the things to optimize: - Cache LLM responses for repeated input shapes. Often 30-50% of re-rank calls have near-duplicate input. - Embed less often. If a user’s signal hasn’t shifted meaningfully, skip the daily embedding refresh. - Move re-rank to a cheaper model. Haiku → smaller open-weight model on owned infrastructure. Each optimization is a 10-30% cost win. Stack them. ## What dominates Across all three scales: 1. LLM tokens (~95%) - the bottleneck. 2. Embedding generation (~3%) - small but real. 3. Database (~2%) - almost free. Three implications: - Don’t optimize the database first. It’s not your bottleneck. A correct database with a wrong LLM is expensive; the reverse is fine. - Pick the smallest model that works. Haiku vs Sonnet is a 4x cost delta on the dominant line item. Test it. - Cache aggressively at the LLM boundary. Even simple input-hash caches save 20-30%. ## What changes the model The numbers above assume one application pattern. Reality varies: - Heavier reasoning (multi-step agent loops): LLM cost goes up 5-10x. Database stays the same. The ratio gets worse. - Lighter generation (single classification, no re-rank): LLM cost drops 90%. Database becomes 20-30% of the bill. Database choices matter more. - Higher write volume (writing every tool call): Database I/O scales meaningfully. OriginChainDB’s commit-window throughput shows up here. Always model your own application before generalizing. ## The OriginChainDB angle This post isn’t a sales pitch - the database is 2% of the bill. But the atomic multi-shape write in OriginChainDB does eliminate one cost line that other architectures pay: the dual-write infrastructure (streams, sync workers, reconciliation jobs) between a primary DB and a vector DB. That’s typically a $200-2000/month line item depending on volume - modest, but it’s real engineering time too. For most teams, the right reason to pick OriginChainDB isn’t cost - it’s the architectural simplicity of one substrate. ## FAQ ### What’s the per-user cost of a typical AI feature? $0.01 to $0.05 per MAU per month for a moderate-complexity feature. Higher for heavier reasoning, lower for light classification. ### Why is the LLM 95% of the cost? Because tokens are individually cheap but you generate billions of them. Even at $0.0001 per call, 1B calls/month is $100K. Storage and database I/O are sub-pennies that don’t add up to the same total. ### Should I optimize my database to reduce AI costs? Probably not. The database is rarely the bottleneck. Optimize the LLM line first. ### What about embedding model costs? Embeddings are <5% of the total bill at moderate scale. Worth caching, not worth obsessing over. ### How does OriginChainDB pricing scale at large volumes? We tier by compute + storage plan. Talk to us - we’ll model your specific workload. ## What to read next - [RAG at production scale: latency budget](https://originchaindb.com/blogs/rag-latency-budget) - the runtime side of the same equation. - [Why we don’t need a separate vector database](https://originchaindb.com/blogs/no-separate-vector-database) - the architectural simplification that saves the dual-write line item. - [Per-key TTL for ephemeral AI data](https://originchaindb.com/blogs/per-key-ttl-agent-memory) - TTL’d caches as a cost-reduction primitive. --- # Automatic Idempotency-Key in the SDK - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/auto-idempotency-key-sdk Published: 2026-05-12T16:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal sdkdxidempotencyapi-design # Automatic Idempotency-Key in the SDK Every mutating SDK call now ships an Idempotency-Key UUID you never type. The bounded-cache check that had to land first, and how the dedupe path works. OriginChainDB May 12, 2026 About 8 min read SDK design / Retries ## A retry needs the same request identity. 1. SDK call Create a request key 2. First attempt Send key and payload 3. Retry Reuse the request key 4. Result Follow the dedupe contract An idempotency key connects attempts for the same operation. Its scope, cache lifetime and endpoint behavior determine what can be safely deduplicated. Every mutating call from the OriginChainDB TypeScript, Python and Go SDKs now attaches an Idempotency-Key UUIDv4 automatically, so a retried write no longer duplicates a row. You pass a key yourself only when you need the same one across processes. The engine caches each key’s response for 24 hours and evicts LRU at 10,000 entries, which is why a fresh key per call cannot grow the cache without bound. A customer pasted a curl from our quickstart into Postman and noticed this: ```plaintext -H "Idempotency-Key: 7c4f...e1" \ ``` His question: “do we require an idempotency key on every write? why is this in the docs?” The honest answer was “the server doesn’t require it; we recommend it; the example hardcoded a placeholder that’s misleading because it makes the header look like a magic constant tied to that endpoint.” A worse answer would have been “yes, you need to mint a fresh UUID before every write.” Neither is the answer he should have to think about. The interesting part isn’t the SDK change - it’s the prerequisite we had to check before flipping the default. ## The header, briefly OriginChainDB’s engine supports the same `Idempotency-Key` header that DynamoDB / Stripe / OpenAI use: - Client picks a string per logical operation (typically a UUID). - Engine records `(client_id, idempotency_key) → response` in an in-memory cache with a 24-hour TTL. - Retries of the same key replay the cached response instead of re-applying the write. A 502-and-retry no longer means a duplicate row. It’s the standard “at-least-once delivery with exactly-once effect” pattern. The engine implements it; the docs document it; the SDKs provided an `idempotency_key` parameter that defaulted to `None`. That `None` default was the bug. The user had to know about idempotency, decide whether to use it, generate a UUID, pass it in. For an SDK aimed at customers building AI features - most of whom have opinions about agents and embeddings, not about HTTP retry semantics - that’s the wrong layering. The fix is one-line per SDK: when the caller passes `None`, generate a UUIDv4 internally and put it on the header. ## The prerequisite Before flipping the default, there’s a question that has to be answered correctly. If you skip this question and ship the new default, you might OOM the engine. The engine’s idempotency cache holds `(client_id, key) → response` for 24 hours. Today, very few callers send the header - most writes have no key, and the cache is sparse. When every SDK call starts sending a fresh UUID, the cache fills up. A tenant writing 1000 rows/second × 86,400 seconds/day = 86.4 million entries in the cache at steady state. At a few hundred bytes each, that’s ~25 GB. We have 2 GB RAM on the smallest tier. The question is: is the cache bounded by anything other than the 24-hour TTL? We had no idea off-hand. We knew the cache existed; we didn’t know the cap. So before any SDK change, we read the cache implementation in the engine: ```rust pub struct IdempotencyCache { inner: Mutex, ttl: Duration, cap: usize, } struct Inner { map: HashMap, order: VecDeque, // LRU order } impl IdempotencyCache { pub fn put(&self, client_id: String, key: String, value: Value) { let mut g = self.inner.lock.unwrap; let ck = CacheKey { client_id, key }; if let Some((_, ts)) = g.map.get(&ck) { // refresh existing - update timestamp + move to back ... } else { // evict oldest if at capacity while g.map.len >= self.cap { if let Some(oldest) = g.order.pop_front { g.map.remove(&oldest); } } g.map.insert(ck.clone, (value, Instant::now)); g.order.push_back(ck); } } ... } ``` LRU-bounded at `cap` (default 10,000 entries) in addition to the 24-hour TTL prune. The cache cannot grow past 10,000 entries. That’s the answer we needed. The “every SDK call sends a fresh UUID could OOM the engine” concern turns out to be moot. The cache will LRU-evict the oldest entries when full. A pathological tenant doing 1 M writes/second would have its cache turn over in 10 ms; idempotency becomes effectively useless at that point, but the engine doesn’t fall over. ## The SDK changes With safety established, the actual changes are minimal: ### TypeScript (`sdk/typescript/src/client.ts`) A `newIdempotencyKey` helper that prefers `crypto.randomUUID` and falls back to `crypto.getRandomValues` for older browser targets. The `_request` plumbing auto-stamps it on mutating methods only (POST / PUT / PATCH / DELETE): ```ts const MUTATING_METHODS = new Set(["POST", "PUT", "PATCH", "DELETE"]); const method = (init.method ?? "GET").toUpperCase; if (MUTATING_METHODS.has(method) && !headers["idempotency-key"]) { headers["idempotency-key"] = newIdempotencyKey; } ``` Caller-provided headers win. If you pass `init.headers["idempotency-key"] = "my-stable-key"`, that’s used; the auto-gen only kicks in when absent. ### Python (`sdk/python/originchain/client.py` and `async_client.py`) Same shape. `_request` auto-stamps. The interesting one is `put_batch`: previously, when the caller passed no key, no header went out for any chunk. With auto-gen, naive code would generate a fresh UUID per chunk, breaking partial-retry dedup: if the third chunk of a 100-chunk batch fails and the caller retries the whole call, the new run gets new per-chunk UUIDs and the engine can’t detect overlap. Fix: mint ONE base UUID at the top of `put_batch`, then derive per-chunk keys deterministically: ```python def put_batch(self, schema, rows, *, idempotency_key=None, chunk=1000): base_idem = idempotency_key or _new_idempotency_key ... for i, r in enumerate(rows): chunk_buf.append(r) if len(chunk_buf) >= chunk: total += self._send_batch(schema, chunk_buf,..., base_idem, i // chunk) ... def _send_batch(self, schema, rows,..., base_idem, chunk_no): headers = {"Idempotency-Key": f"{base_idem}-c{chunk_no}"} ... ``` A retry of the SAME logical `put_batch` call (same process, same batch object) gets the same base UUID - chunks dedup. Callers who want cross-process retry of the same logical action pass `idempotency_key ="abc"` explicitly; chunks become `"abc-c0"`, `"abc-c1"`, dedupable forever. ### Go (`sdk/go/client.go`) Stdlib `crypto/rand` + `encoding/hex`, zero new dependencies (the Go SDK’s `go.sum` was empty and we wanted to keep it that way): ```go func newIdempotencyKey string { var b [16]byte if _, err := rand.Read(b[:]); err != nil { return "" } b[6] = (b[6] & 0x0f) | 0x40 // version 4 b[8] = (b[8] & 0x3f) | 0x80 // variant 1 out := make([]byte, 36) hex.Encode(out[0:8], b[0:4]) out[8] = '-' hex.Encode(out[9:13], b[4:6]) out[13] = '-' hex.Encode(out[14:18], b[6:8]) out[18] = '-' hex.Encode(out[19:23], b[8:10]) out[23] = '-' hex.Encode(out[24:36], b[10:16]) return string(out) } func isMutatingMethod(method string) bool { switch strings.ToUpper(method) { case http.MethodPost, http.MethodPut, http.MethodPatch, http.MethodDelete: return true } return false } ``` The shared `request` function auto-attaches the header for mutating methods only. GETs never consume an idempotency cache slot. ## Docs cleanup The hardcoded `Idempotency-Key: 7c4f...e1` in the quickstart was worse than useless - it implied the header was required and that the value was constant. Both wrong. Three docs pages got the header dropped from the minimal curl examples: - `quickstart.astro` - `insertCurl` no longer carries the header. The prose now reads: “Retries are safe by default: the SDKs auto-attach an Idempotency-Key on every mutating call. Set your own only when hitting raw HTTP and you need cross-process retry semantics.” - `insert.astro` - same treatment on the single-row and batch curl examples. - `sdk.astro` - the Python `rows.put` example no longer passes `idempotency_key=str(uuid.uuid4)` (the SDK does it). The API reference page `api.astro` keeps the header in the optional- headers table - that’s the right place for it. We didn’t delete the concept, just stopped requiring users to think about it. ## Tests Across the three SDKs, 15 new test assertions: - Auto-generated key is canonical UUIDv4 (hex32 or hyphenated 36 depending on language). - Caller-supplied key wins (override still works). - GET never sends Idempotency-Key. - Python `put_batch` per-chunk keys share a base UUID - a 3-chunk batch produces `{base}-c0`, `{base}-c1`, `{base}-c2` with one common prefix. 10 TS tests + 34 Python tests + 3 Go tests pass. ## Honest scope - The header is still optional on the engine. This change is pure SDK-side. A raw `curl` user who doesn’t send the header gets the same behaviour as before (no dedup); the engine doesn’t require it. We just made the SDK helpful by default. - The cache bound is process-local. Each tenant has one cache today. With multi-writer (on the roadmap), the cache would need to be replicated or scoped per writer. Cross-replica idempotency is a v2 problem; we’ll keep this property in mind. - The account/management API doesn’t dedupe. The SDK auto-stamps on calls to both the engine and the management API, but management endpoints (signup, signin, instance-create) don’t consume the header today. The header arrives, is ignored, future consumption is a one-line opt-in if we ever want it. ## What this is A small DX win, but the kind that compounds. Customers were typing UUIDs to defend against a failure mode they didn’t have to think about. The right shape is “do it for them automatically; let them override when they need to.” That’s the default DynamoDB / Stripe / OpenAI established for transactional APIs; we should match it. If we’d shipped the SDK change first and discovered the cache was unbounded only in production, we would have OOM’d a fleet of tenants under a write-heavy workload. ## Try it Update the SDK: ```bash npm install --save @originchain/sdk@latest pip install --upgrade originchain go get -u github.com/originchain-ai/originchain-go ``` Make a write. Don’t pass an idempotency key. Retry on failure. It deduplicates. You didn’t have to think about it. One bearer. One atomic write. One header you no longer have to generate. ## What to read next - [Idempotent tool calls](https://originchaindb.com/blogs/idempotent-tool-calls) - the agent-loop pattern this header exists for. - [Backpressure: 429 + Retry-After](https://originchaindb.com/blogs/backpressure-429) - when the engine asks you to retry, and how to do it politely. --- # Write backpressure: 429 + Retry-After - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/backpressure-429 Published: 2026-05-06T10:26:12.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-14T15:11:01.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal backpressurerate-limitingdesignreliabilityapi # Write backpressure: 429 + Retry-After OriginChainDB refuses overload with HTTP 429 and a Retry-After computed from the token bucket's refill schedule. The per-API-key limits, and retry code. OriginChainDB Team May 6, 2026 About 5 min read API design / Backpressure ## Make room before the next attempt. 1. Request Submit the operation 2. HTTP 429 Read the retry guidance 3. Wait Back off with a bounded policy 4. Try again Preserve request identity Backpressure is a signal to the client. A bounded retry policy respects server guidance instead of immediately repeating an overloaded request. TL;DR - When a database can’t keep up with incoming writes, it has two choices: refuse them with a 429 + Retry-After (graceful) or accept them and silently fail later (catastrophic). OriginChainDB’s backpressure is per-API-key, expressed as 429 with a precise Retry-After, and tuned so a polite client recovers within seconds. The bucket is checked at the edge, so the 429 fires within microseconds rather than stalling the socket. ## What backpressure is for Every database has a finite ceiling. Writes per second, concurrent connections, durable-commit rate - pick your bottleneck. When you hit it, the database has options: 1. Queue indefinitely. Latencies spike, eventually clients time out, you find out from a pager. 2. Drop writes silently. Clients think their data is durable; it isn’t. 3. Refuse cleanly with a clear protocol contract. The client knows it failed, knows when to retry, and the database recovers without anyone needing to intervene. OriginChainDB picks option 3. The protocol is HTTP 429 (Too Many Requests) with a `Retry-After` header expressing the wait in seconds. ## The per-key model Limits are scoped per API key, not per tenant or per endpoint. Each key has its own bucket; one runaway script can’t starve the others. The buckets are: - Sustained write rate - a token-bucket that refills at a configured RPS and absorbs bursts up to a small multiple. - Sustained read rate - same model, separately tracked. - Concurrent in-flight requests - capped at a small number per key (default 64). Burst-y clients don’t fork-bomb. When a request hits a depleted bucket, the response is: ```plaintext HTTP/1.1 429 Too Many Requests Retry-After: 2 Content-Type: application/json {"error":{"code":"rate_limited","message":"write rate exceeded; retry after 2s","key_id":"k_abc..."}} ``` `Retry-After` is computed from the bucket’s refill schedule, not from a fixed-back-off table. If the server expects to drain in 1.7 seconds it returns `2`; if it expects 0.3 seconds it returns `1`. The client gets a real estimate, not a wild guess. ## Why per-key, not per-tenant Tenants on OriginChainDB’s dedicated configurations map to dedicated instances, and each awake Free database runs in its own engine process; the substrate doesn’t multi-tenant inside a single process. So per-tenant limits are enforced at the deployment level, not the API level. What the API key bucket protects against is your own application misbehaving - a runaway worker, a debugging script that forgot to back off, an autonomous agent that’s stuck in a tool-call loop. The key system gives you a way to provision multiple keys for a single tenant (one per service) and isolate them. ## Why HTTP 429, specifically A few reasons: - It’s the right semantic. 429 is “you’re sending too many requests.” 503 (“service unavailable”) would be wrong - the database isn’t down, just busy on your behalf. 500 would be even worse - that means we screwed up. - HTTP libraries handle it natively. Most HTTP clients have built-in 429 retry with `Retry-After` honoring. You usually don’t even write retry logic. - It’s auditable. A 429 in your monitoring tells you exactly which key is over budget; a 503 doesn’t distinguish between “rate-limited” and “everything’s broken.” ## The client side What a polite client does: ```typescript async function writeWithBackoff(payload: any, attempt = 0): Promise { const res = await fetch("/api/write", { method: "POST", body: JSON.stringify(payload), }); if (res.status === 429) { if (attempt >= 5) throw new Error("rate limit exceeded retry budget"); const delay = parseInt(res.headers.get("retry-after") ?? "1", 10); await sleep(delay * 1000); return writeWithBackoff(payload, attempt + 1); } if (!res.ok) throw new Error(`http ${res.status}`); } ``` Three things to notice: 1. Honor `Retry-After` - it’s not an arbitrary number, it’s the substrate’s actual refill estimate. Polling sooner just generates more 429s. 2. Cap the retry attempts - if you’re hitting 429 on attempt 6, the client is wrong, not the substrate. 3. No exponential backoff needed - `Retry-After` already encodes the right delay. Adding jitter on top is fine but not required. What an impolite client looks like: tight retry loop, ignores `Retry-After`, gets stuck. The substrate handles this gracefully - it’ll keep returning 429s - but the client wastes its own runtime polling for no reason. ## Limits at a glance Default values for an entry configuration: | Bucket | Sustained | Burst | Notes | | --- | --- | --- | --- | | Writes | 1,000/sec | 4,000 | Per API key | | Reads | 10,000/sec | 30,000 | Per API key | | Concurrent in-flight | - | 64 | Hard cap | | Vector search | 200/sec | 1,000 | Counted separately from reads | Higher tiers raise sustained and burst proportionally. Enterprise plans have configurable per-key limits. These numbers are not a marketing pitch - they’re conservative enough that a well-behaved client will rarely see them, and aggressive enough that the bucket actually does work when something’s wrong. ## What the substrate does at the edge When a bucket is exhausted, the substrate doesn’t queue the request indefinitely. The 429 fires within microseconds - there’s no socket-level “wait until I have capacity” stall. This is deliberate. A queue at the edge means latency variance the client can’t see; an immediate 429 means the client knows immediately and can decide what to do. For autonomous agent workloads in particular, knowing “I was rate-limited” is far more useful than seeing every request take an unpredictable amount of time. ## FAQ ### What is HTTP 429? HTTP 429 (Too Many Requests) is the standard status code for “you’re being rate-limited.” It’s accompanied by a `Retry-After` header indicating how long to wait before the client should try again. ### Will my agent loop survive being rate-limited? Yes, as long as it honors `Retry-After`. Most HTTP libraries do this automatically. If you wrote a custom client, make sure the retry path reads the header and sleeps the indicated amount. ### Can I raise my limits? Yes - enterprise plans let you configure per-key limits explicitly. For the free database and starter configurations, the default is what you get. ### Are limits per-IP or per-API-key? Per-API-key. IP-level limits would penalize legitimate traffic from corporate NATs. ### What about read-after-write consistency under backpressure? Unaffected. Backpressure refuses writes that would overrun the substrate; it doesn’t change consistency semantics for accepted writes. ## What to read next - [Idempotent tool calls](https://originchaindb.com/blogs/idempotent-tool-calls) - pairs naturally with backpressure (retry-with-key-dedup). - [RAG latency budget](https://originchaindb.com/blogs/rag-latency-budget) - how 429s show up (or don’t) in your end-to-end latency. --- # Our depth-first roadmap to 1.0 - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/depth-first-roadmap-to-1-0 Published: 2026-05-04T20:59:15.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal roadmapproductdesignarchitecture # Our depth-first roadmap to 1.0 The roadmap to 1.0 is HA, fuzzing, optimiser, EXPLAIN, multi-writer, online schema change - in that order. What that buys, and what it leaves off. OriginChainDB Team May 4, 2026 About 7 min read Product engineering / Roadmap ## Build depth into each layer. 1. Reliability Failure handling and replication 2. Verification Fuzzing and regression tests 3. Query planning Optimizer and EXPLAIN 4. Evolution Topology and schema changes The article records a prioritization strategy. These layers illustrate that strategy, rather than a live release schedule or availability checklist. TL;DR - OriginChainDB’s roadmap to 1.0 is depth-first by deliberate choice. Six months of HA → fuzzing → optimiser → EXPLAIN → multi-writer → online-schema, and zero new “shapes” of data (no graph, no full-text, no time-series specialty types) in that window. No step starts until the previous one has run in production for at least one paying customer. ## The breadth temptation When you’re building an AI-native database, the easy roadmap writes itself. Customers ask: - “Can you add a graph engine for relationship queries?” - “Can you add full-text search so we don’t need OpenSearch?” - “Can you add time-series compression for telemetry workloads?” - “Can you add geospatial indexes?” Every one of those asks is reasonable. Every one of them is also worth a year of engineering. If you say yes to all of them, you ship a database that does eight things mediocrely. The vector substrate handles small loads but falls over at scale. The graph engine works on toy datasets but doesn’t have a real query optimiser. The full-text search is a port of an old library and has half the recall of a dedicated FTS engine. Customers who pick you because of the “it does everything” pitch end up dropping you because of the “but none of it works in production” reality. We’ve watched that movie before. We’re not making it again. ## What depth-first means here Depth-first means: every quarter, we ship one capability that fully works at scale, with the boring parts (HA, observability, recovery, performance under load) done before we declare it shippable. We don’t move on until that’s true. Concretely, the locked plan to 1.0: 1. HA + automatic failover - done. Snapshot bootstrap landed; chaos drill passes. To be precise about what that buys: a write is acknowledged only once it is flushed to durable storage on the instance, and a standby can be promoted to take writes, with automatic promotion available as an opt-in that refuses a standby which is not fully caught up. Replication to that standby is asynchronous, so an abrupt loss of the primary can still cost the most recent writes - the seeding gap is closed, not the replication gap. ([Read the post](https://originchaindb.com/blogs/ha-snapshot-bootstrap).) 2. Crash-testing the substrate - in progress. Crash-test the substrate at every durability boundary; build automation that finds invariant violations under fault injection. (Read the post.) 3. Cost-based optimiser - next. Replace the current rule-based plan with one that estimates I/O cost across shapes and picks the cheapest path. 4. EXPLAIN - alongside (3). Show developers what the planner actually did so debugging slow queries is a 30-second loop, not a four-hour mystery. 5. Multi-writer - after (4). Allow more than one concurrent writer without breaking the per-key consistency story. This is a hard problem; we want the rest of the substrate proven before we touch it. 6. Online schema evolution - last. Shape upgrades that don’t require any read or write pause, with the dual-read transform automated end-to-end. Each step is one full quarter of engineering. Some are longer. None overlap; we don’t start step N+1 until step N is in production for at least one paying customer. ## What we’re explicitly not doing in this window | Feature | Status | Why not yet | | --- | --- | --- | | Vector search at >100M | Works at <10M today | Optimiser work in step 3 unblocks larger ANN | | Graph traversal | Not on roadmap | Adjacency-list shapes can model it; no native engine before 1.0 | | Full-text search | Not on roadmap | Indexed string equality covers most AI use cases; FTS post-1.0 | | Time-series compression | Not on roadmap | Standard shapes work; columnar compression later | | Geospatial indexes | Not on roadmap | Niche enough we wait for a paying customer to ask | | SQL surface | Not on roadmap | We have a typed query API; SQL adds a parsing layer with no functional benefit | | Self-hosted distribution | Not on roadmap | Managed-cloud only - full stop | These will come, eventually, in some order. Probably not in the next six months. We’d rather you can rely on the things we do ship than have a long list of things that kind of work. ## Why this is the right bet for AI-native Three reasons. 1. AI workloads punish weak engines harder than transactional workloads do. A traditional CRUD app touches a few rows per request. An AI agent loops, hammering hundreds of writes and reads per session, often with retries and tool-call branches. A database that’s “fine for most workloads” turns into a bottleneck on hour two of an agent run. The substrate has to be tight before the layer above it matters. 2. The features customers actually need first are infrastructure, not data shapes. Talk to teams running AI features in production. Their top complaints are: “the latency variance ruined our agent loop,” “the failover lost data and we found out from a customer,” “we can’t tell why this query is slow,” “schema migration broke the eval pipeline.” These are all infrastructure problems, not data-shape problems. Depth-first addresses them; breadth-first doesn’t. 3. AI-native means less ceremony, not more shapes. Adding a graph engine doesn’t reduce the developer’s mental load - it adds a new query language to learn, a new failure mode to debug, a new index to size. Key shapes already model graphs (adjacency-list shape + index on the “to” side) without inventing new vocabulary. The right answer for “AI-native” is fewer specialty engines, not more. ## When does this stop? When 1.0 ships, the substrate is verified at production scale, the planner is honest about cost, multi-writer is online, and schema evolution is invisible. Then we start saying yes to the breadth requests - and even then, only for the ones a paying customer is asking for. The roadmap after 1.0 isn’t pre-decided. It’ll be customer-driven, and we’d rather have a small list of customers asking for the same thing than a long list of features no one is paying for. ## What this means for you, today If you’re evaluating OriginChainDB right now, the honest answer is: - Use us if your workload is JSON records + vectors + light secondary indexes, and you care more about p99 latency, atomic multi-shape writes, and durable-on-commit writes with automatic failover than about feature breadth. - Wait if you need full-text search, geospatial, graph traversal, or columnar analytics. None of those land before 1.0. - Talk to us if you’re somewhere in between. Sometimes a workload that seems to need a specialty engine actually models cleanly as shapes. The depth-first decision is a commitment to the customers we have, not the features we could ship. We’d rather lose a deal because we don’t have feature X than win it and miss the SLA we promised. ## FAQ ### Why isn’t vector search higher on the priority list? Vector search at small-to-medium scale works today. The ANN performance we have is good enough for sub-10M-vector corpuses. The optimiser work in step 3 of the roadmap is what unblocks much-larger ANN; until then, the bottleneck is plan selection, not the index. ### When will graph queries / full-text / time-series land? Post-1.0, customer-driven. We won’t pre-commit to a quarter for any of them - that’s the point of depth-first. If you need one of these now, we’re not the right database. ### What does 1.0 mean concretely? A version of the substrate where every step in the depth-first roadmap is shipped, the cost-based optimiser is verified against real workloads, multi-writer is GA, and online schema evolution requires zero application changes. The number on the version isn’t the point - those properties are. ### Are you taking enterprise customers in this window? Yes, but with eyes open. We’re transparent about what works (depth-first capabilities above) and what doesn’t yet (breadth). Enterprise customers in this window are the ones aligned on “we want the substrate done well first.” ### Will the roadmap change? The order might shift if a step takes longer than expected. The depth-first principle won’t. Adding a new top-level capability before 1.0 is the change we won’t make. ## What to read next - [Building an AI-native database](https://originchaindb.com/architecture) - the architectural premise. - [HA snapshot bootstrap](https://originchaindb.com/blogs/ha-snapshot-bootstrap) - depth-first step 1, in production. - How we crash-test the substrate - depth-first step 2, in progress. --- # Eight comparison pages in one afternoon - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/eight-vs-pages-three-agents-one-afternoon Published: 2026-05-04T16:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal engineeringai-toolsagentsmarketing-site # Eight comparison pages in one afternoon Eight competitor comparison pages shipped in an afternoon by three coding agents run in parallel. The prompt, the fairness rubric, and what agents can't do. OriginChainDB May 4, 2026 About 7 min read Engineering workflow / Parallel work ## Parallel drafts. One review standard. 1. Scope One bounded page per assignment 2. Research Keep a common evidence bar 3. Write Use an agreed page structure 4. Review Check claims, links and layout Parallel work starts with explicit ownership and ends with a shared review. Research quality and editorial judgment remain human responsibilities. Three coding agents running in parallel on separate worktrees produced eight competitor comparison pages in about three hours of wall clock, prompt to merged PR. What made the output keepable was not the agents: it was a structural template, a batch of two or three pages per agent, and a fairness rubric, all written before any agent started. Comparison pages are a strange genre. Customers want them - “how does OriginChainDB compare to Pinecone?” is the most common question on a sales call after pricing. Marketing pages want them - they rank for high-intent queries like “pinecone vs…” and “alternatives to…”. And yet, building eight of them by hand takes a week of writer time, an editor’s review pass, and three rounds of “is this fair?” pushback. We shipped eight in under an afternoon. `/vs/postgres`, `/vs/pinecone`, `/vs/weaviate`, `/vs/qdrant`, `/vs/milvus`, `/vs/supabase`, `/vs/neon`, `/vs/mongodb` - each a standalone Astro page with a comparison table, an honest “what they do well” section, an “how OriginChainDB differs” section, a code example contrasting the two stacks, and a closing call-to-action. Each agent worked on its own worktree with a small batch of pages. ## The pattern: structural template + small batches + fairness rubric Three decisions made the difference between “agents wrote a content pile we threw away” and “agents shipped pages we kept.” ### Decision 1: Pre-write a strict structural template The biggest failure mode for AI-written marketing pages is structural drift. Page 1 leads with a code example; page 2 leads with a comparison table; page 3 leads with a quote. Inconsistency reads as carelessness, and worse, it makes the pages feel like they were written by an LLM - which is the one signal we needed to avoid. The fix was a structural template, written by a human, that every page must match 1:1: ```plaintext 1. Hero (one sentence, what they are, one sentence, what we are) 2. Comparison table (8–12 rows, identical row labels across all pages) 3. "What [competitor] does well" (3 paragraphs, no hedging) 4. "How OriginChain differs" (3 paragraphs, concrete, no hype) 5. Code example (their stack vs ours, side-by-side) 6. "When to pick which" (decision rubric, 3 bullets) 7. CTA (links to /docs/quickstart and /pricing) ``` The row labels in the comparison table are fixed: Atomic across shapes, SQL with JOINs, HNSW recall@10 at 100k, BM25 full-text, Graph traversal, Tenant isolation, Replication and failover, Pricing model, Self-host story. Same labels across all eight pages. Readers comparing two of our pages immediately see the same axes - that is the whole reason a comparison table is worth printing. ### Decision 2: Small batches per agent Each agent got 2–3 pages, not 8. This sounds like a small detail. It is not. An agent given 8 pages will compress. By page 4 the descriptions get shorter, the code examples get more generic, and the “what they do well” paragraphs start sounding suspiciously similar across competitors. Context budget runs out, and quality silently drops. An agent given 2–3 pages keeps the working set small, the writing fresh, and the per-page time roughly constant. We split 8 pages across 3 agents (3 + 3 + 2) and saw uniform quality across all eight outputs. The cost of running three sessions is a tiny fraction of the cost of having a human re-edit a degraded batch. ### Decision 3: Fairness as a hard rubric The most important sentence in the prompt was this one: Be FAIR. Pinecone is a real product used by real teams. Explain what they do well - managed vector search, fast time-to-first-query, mature SDKs, dense ecosystem of integrations. Then explain how OriginChain differs. Do not be hostile, do not strawman, do not omit a competitor’s strengths to make ours look better. Hostile comparisons rank poorly on Google, read poorly to engineers, and damage the brand they’re trying to elevate. Fairness is not a soft preference. It is a hard rubric, because: - Hostile comparisons get flagged by Google’s helpful-content classifiers and rank below neutral writeups. - Engineers comparing tools can smell hype in two paragraphs and bounce. - The pages need to survive being read by the competitor’s engineers without making us look unserious. The agents complied with the rubric cleanly. The Pinecone page calls out Pinecone’s managed-service polish and SDK quality before pivoting to atomicity-across-shapes. The Postgres page calls out pgvector and the Postgres ecosystem before pivoting to native multi-shape. The Neo4j… we didn’t do Neo4j, that one’s still on the list. ## A sample agent prompt (excerpt) This is the actual middle of one of the prompts we used: ```plaintext You are writing /vs/pinecone.astro for the OriginChain marketing site. Follow the structural template in /docs/internal/vs-template.md exactly. Do not deviate from row order in the comparison table. Do not invent performance numbers. The only numbers you may cite for OriginChain are: - HNSW fast: recall@10 = 0.69, p99 = 37 ms (100k vectors, D=128) - HNSW high_recall: recall@10 = 0.96, p99 = 109 ms - Asynchronous standby replication with fenced automatic failover - crash testing at 4 durability boundaries, 5,000 iters/boundary nightly For Pinecone numbers, link to their public docs; do not state internal benchmarks you have not run. If a row in the table is genuinely "both support this," mark it Yes / Yes. Do not gaslight rows green for us that are not green for us. The "what Pinecone does well" section must be 3 paragraphs. Suggested angles: managed-service quality, SDK ergonomics, integration ecosystem. The "how OriginChain differs" section must be 3 paragraphs. Suggested angles: atomicity across rows + vectors + FTS + graph in one write, single bearer, sub-millisecond SQL alongside vector. Do not use the words "revolutionize", "unleash", "game-changer", "next-generation", "paradigm", or "empower." ``` The word ban looks petty. It is not. Those six words are the dominant tells of LLM marketing prose, and banning them forces the agent into more concrete language. Ban one and another grows back; ban six and the prose tightens. ## What agents don’t replace This is the part that matters, and the part most “agents wrote our marketing site” stories skip. Agents do not replace editorial judgment. Every decision the agents made was already encoded in the prompt: - The structural template (1 hour to write, used 8 times). - The fixed row labels for the comparison table (45 min, including arguing about which 9 axes to settle on). - The fairness rubric (15 min, but every word matters). - The numbers list (we have to be ready to cite these on a sales call). - The word ban (5 min, but it took six months of reading bad LLM prose to know which words to ban). That is roughly 80% of the actual creative work. The agent contribution is the last 20% - turning the rubric into eight specific pages of prose, each calibrated to a specific competitor. This is the part that most teams misjudge in both directions. The bullish version is “agents will write our marketing site for us.” The bearish version is “AI-generated content is slop and we won’t use it.” The realistic version is “agents are an extremely fast last-mile, and the editorial work that decides whether the output is good is exactly the work humans were always going to do.” ## What this is NOT A few honest scope notes: - There is no /vs/ index page yet. Each page is reachable via direct URL - `/vs/postgres`, `/vs/pinecone`, etc. - but there is no landing page that lists all eight. The index is on the punch-list; for now, link to the individual pages. - The pages are static. No A/B tests, no dynamic copy based on referrer, no conversion-tracking instrumentation beyond the standard analytics. Comparison pages tend to convert on long tail SEO, not paid acquisition, so the simplicity is intentional. - We did not run benchmarks against the competitors. Where the table cites a competitor capability, it is sourced from their public documentation. We cite our own numbers (HNSW recall@10, p99, crash-injection iters) from our own benchmarks. If a competitor publishes new numbers, the page can be regenerated against the updated source in one agent run. - Eight pages is not the universe. Cassandra, Redis, ClickHouse, DynamoDB, Couchbase, Elasticsearch - there are at least a dozen more comparisons that make sense. We shipped eight because eight is the batch where we know our positioning is sharp; the next batch lands when the next batch’s positioning is sharp. ## Try it If you are evaluating OriginChainDB against another database, the eight comparison pages are at: - [/vs/postgres](https://originchaindb.com/vs/postgres) - [/vs/pinecone](https://originchaindb.com/vs/pinecone) - [/vs/weaviate](https://originchaindb.com/vs/weaviate) - [/vs/qdrant](https://originchaindb.com/vs/qdrant) - [/vs/milvus](https://originchaindb.com/vs/milvus) - [/vs/supabase](https://originchaindb.com/vs/supabase) - [/vs/neon](https://originchaindb.com/vs/neon) - [/vs/mongodb](https://originchaindb.com/vs/mongodb) Each is a fair, technical writeup - not a hit piece. If the comparison you want isn’t there yet, the [docs](https://originchaindb.com/docs) cover the substrate-level shape of OriginChainDB in detail; the comparison page is just the last-mile rendering against a specific alternative. Three agents. One afternoon. Eight pages. The pattern works. --- # From a 20-second dashboard to 300 ms - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/from-20s-dashboard-to-300ms Published: 2026-05-12T20:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal infrastructureperformanceplatformoperations # From a 20-second dashboard to 300 ms Reading a tenant credential over a remote command cost 4-15 seconds per dashboard page. A managed parameter store cut it to ~150 ms. The migration path. OriginChainDB May 12, 2026 About 7 min read Performance / Request path ## Remove remote work from the hot path. 1. Remote command Resolve credentials per request 2. Dashboard read Wait for the dependency 3. Parameter store Use the configured source 4. Local cache Reuse a bounded result The case study compares credential lookup paths and caching. Timing in the article belongs to its measured setup, not a universal latency promise. The dashboard’s cold-path latency was a credential fetch, not the database. Reading each tenant’s bearer token with a remote command to the instance cost 4-15 seconds; moving that read to a managed parameter store cut it to ~150 ms, and the cold page load from 20 seconds to ~300 ms. A customer mentioned in a thread that the schema-list page on their OriginChainDB dashboard took “like around 20 seconds.” Not a feature request: a complaint. The page renders four schemas. Each is a hundred-byte manifest. The engine serving them, behind a TLS connection, responds to `GET /v1/tenants/.../schemas` in ~350 ms end-to-end including the TLS handshake. So where do the other ~19 seconds go? They went into how the dashboard backend authenticated to the tenant engine on the user’s behalf. ## The architecture and the problem OriginChainDB’s customer dashboard lives at `app.originchain.ai`. The backend serving that dashboard does cookie-auth against the user. When the user opens a schema page, the backend translates the request into a tenant-bearer call against the customer’s dedicated engine instance. The tenant bearer is the customer’s secret; it is written onto the instance at provision time. The backend needs that bearer to make the proxy call. It doesn’t have it in memory. It has to ask the tenant box for it. The original implementation issued a remote command to the instance to read the credential file off disk and return its contents, then polled for the result: ```plaintext 1. issue a remote "read this credential" command to the instance 2. wait for the command to be dispatched 3. the instance runs it and returns the credential on stdout 4. poll for completion every 500 ms 5. once complete, read the result back ``` The remote-command API we used is beautifully versatile and operationally awful for a request that’s on the hot path of a page load. The wire flow involves spawning a helper process (~1 s cold start), a multi-second control round-trip to dispatch the command, a hop to the agent on the target instance, and then repeated polling - each poll paying the same cold-start cost again. Cold path: 4–15 s typical, with a tail at 20 s. The result was cached in-process for 5 minutes, so most page loads were fast. But every cache miss - first load of the day, first load after a backend restart, first load after the 5-min idle - paid the full price. When the user hits a dashboard page on a Monday morning, that’s the cache miss. Hence “20 seconds.” ## Why the original design picked a remote command Hindsight makes this look silly, but at the time it was reasonable: - The credential file was the source of truth for the bearer. Other operational code already used the same remote-command mechanism for things like log-tailing and one-shot diagnostic queries. Reusing the pattern felt clean. - The bearer was already on the box (written at boot during provisioning). Storing it somewhere else felt like duplicate state. - Permissions were already set up for the remote-command path. Adding another permission (parameter-store read) felt like work. These are individually fine reasons. The sum of them, weighted by the fact that the call happens on every dashboard request after a 5-minute gap, was a bad trade. ## The fix Three things: ### 1. Write the bearer to a managed parameter store at provision time Per tenant, encrypted at rest, under a per-instance path. The provision flow already mints the bearer and writes it to the instance; adding a parameter-store put alongside is a few lines in the instance-create handler: ```rust let bearer = generate_bearer; let bearer_hash = password::hash(&bearer)?; //... existing DB insert + provisioning trigger... bearer_store::put(instance_id, &bearer).await?; ``` The decrypt permission was already granted to the backend’s role; the put/get permission on the per-tenant path is a fresh statement. ### 2. Replace the bearer fetch with a direct call to the parameter store ```rust pub async fn get(instance_id: Uuid) -> AppResult> { let client = param_client.await?; let name = format!("/oc/tenant-bearers/{}", instance_id); match client.get_parameter .name(&name) .with_decryption(true) .send.await { Ok(out) => Ok(out.parameter.and_then(|p| p.value).map(String::from)), Err(err) if matches!(err.as_service_error, Some(GetParameterError::ParameterNotFound(_))) => Ok(None), Err(e) => Err(AppError::Internal(format!("param store: {e}"))), } } ``` The cold call: one signed HTTPS GET, decrypt server-side, response in ~150 ms. The client is a process-wide singleton so the connection pool warms once and every subsequent call reuses connections. ### 3. Backwards-compatible fallback for unmigrated instances The catch: existing tenants didn’t have their bearer in the parameter store yet. We could have run a one-shot migration script, but the correctness story is cleaner if the backend handles it automatically: ```rust // Try the parameter store first. ~150 ms cold. if let Some(bearer) = bearer_store::get(instance_id).await? { return Ok(cache_and_return(bearer)); } // Fall through to the legacy remote-command read of the credential. let bearer = fetch_tenant_bearer_via_remote_command(ctx).await?; // Opportunistically backfill the parameter store. Next call lands fast. if let Err(e) = bearer_store::put(instance_id, &bearer).await { warn!(error = %e, "bearer parameter-store backfill failed"); } Ok(cache_and_return(bearer)) ``` Three properties: - New tenants always land in the parameter store via the provision flow. Fast path from day one. - Existing tenants pay the legacy slow path ONCE - the very next call is fast. - The backfill is best-effort. If permissions are misconfigured or the store is briefly unavailable, the tenant still works; we just stay on the slow path until the issue is fixed. We also pre-seeded the live demo tenant’s bearer directly so the first dashboard load after the deploy was fast, not 15 seconds. Self-bootstrap, when you can. ## Bumping the in-process cache TTL The legacy slow path was so painful that we’d capped the in-process cache at 5 minutes - the cost of a cache miss was so high that we wanted to refresh aggressively in case the bearer had rotated externally. With the new path at ~150 ms, the cost of a miss is small. We bumped the cache TTL to 1 hour. The bearer rotates only on an explicit `rotate-bearer` call (which invalidates the cache locally) or rare operator intervention; 1 hour catches an out-of-band rotation soon enough without paying needless parameter-store calls. ## Live numbers, end-to-end After deploying: | Path | Before | After | | --- | --- | --- | | Cache miss (1st load / 1h idle) | 4–15 s | ~300 ms | | Cache hit (tab switch) | <1 ms | <1 ms | | Tail (p99 of cache miss) | 20+ s | ~500 ms | The 300 ms vs 150 ms gap is the rest of the dashboard request: DB lookup of the user’s org + instance ownership check, the proxy call to the tenant engine, the JSON serialise. The parameter-store call itself is the smallest part of the new latency budget. p99 dropped from “20 seconds” to “500 ms” - a 40× improvement on the worst case, ~100× on the mean. The hot path was already fast; the fix was about the cold path. ## What this is NOT - A rotation system. The bearer can still rotate via the existing `POST /v1/instances/:id/rotate-bearer` endpoint; this fix doesn’t add automatic rotation on a schedule. That’s a separate compliance feature. - A trust-store. The parameter store uses an account-level encryption key. We didn’t move to per-tenant keys here; the trade-off is “blast radius if the key is compromised” vs “ops complexity at 100 tenants times $1/month.” The blast radius is the same as it was before; the perf is way better. - A free-lunch claim. The parameter store has rate limits. For our load - a few requests per dashboard page, behind an in-process cache - we’ll never hit them. A tenant with 50,000 simultaneous dashboard users on a cache-cold restart would; not our problem today. ## Why this is worth blogging Most database vendor blog posts are about the hot path: the smart index, the magic protocol, the new compression scheme. Cold paths shape product perception just as much. A 20-second first load makes your dashboard feel broken; a 300 ms first load makes it feel snappy even if the rest of the system is the same. If you’re building a SaaS control plane and reaching for a remote-shell call because “it’s already there,” consider the alternative. A managed parameter store is the right primitive for any sub-kilobyte per-tenant secret, and the migration is a fallback plus a backfill. ## Try it Open the schema page on your tenant. It’ll feel faster than it used to. ## What to read next - [RAG latency budget](https://originchaindb.com/blogs/rag-latency-budget) - the same budgeting exercise on the query path. - [HA snapshot bootstrap](https://originchaindb.com/blogs/ha-snapshot-bootstrap) - how a replica gets seeded without losing the writes that land while it seeds. --- # HA snapshot bootstrap: the cutover gap - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/ha-snapshot-bootstrap Published: 2026-05-04T20:59:15.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal hareplicationarchitecturedurabilityfailover # HA snapshot bootstrap: the cutover gap Snapshot-based bootstrap closes the window where writes land while a replica is still seeding. How the cutover works, and what it still won't protect. OriginChainDB Team May 4, 2026 About 6 min read Replication / Bootstrap ## Bridge the snapshot and the live stream. 1. Snapshot Capture a starting state 2. Transfer Seed the new replica 3. Replay Apply changes after the snapshot 4. Cutover Use the promotion procedure Replica bootstrap needs a defined handoff between a snapshot and subsequent writes. Recovery guarantees still depend on replication and promotion behavior. TL;DR - Most database failovers silently drop writes that landed while a new replica was being seeded. OriginChainDB ships a snapshot-based bootstrap: a new replica streams a consistent point-in-time snapshot, then catches up on every write committed after that point, before it is allowed to take over. That closes the seeding gap. It does not make failover lossless - replication to the standby is asynchronous, so an abrupt loss of the primary can still cost the most recent acknowledged writes. Both halves of that are worth saying out loud. ## The problem with naive failover The typical “high-availability” Postgres or MySQL setup looks like this: one writer, one or more streaming replicas, a vote-based failover that promotes a replica when the writer goes silent. It works, mostly. Until it doesn’t. The failure mode that catches everyone is the cutover gap. A new follower joins the cluster. It needs to catch up to the current state. The naive approach is to copy the writer’s data files, then start streaming changes from the point the snapshot was taken at. Sounds reasonable. The problem is that “copying data files while writes are happening” is itself an operation on a moving target, and the point you stamp is always approximate. When that follower gets promoted later, it serves traffic from a state that’s missing the writes that landed in the gap between snapshot start and snapshot end. Customers see “I just placed an order - where is it?” and the team spends three hours diffing replicas to figure out what got dropped. This is the bug we refused to ship. ## What OriginChainDB does instead OriginChainDB’s follower bootstrap is a two-phase exchange between the writer and the joining follower: 1. Snapshot phase. The writer freezes its current commit position (call it `W₀`), takes a consistent snapshot of every shape - rows, vectors, indexes, the lot - and streams it to the follower. The follower stores this as its starting point. 2. Catch-up phase. The follower then asks the writer for every write committed between `W₀` and the writer’s current position, applies them in order, and continues replicating live changes after that. The key invariant: a follower is not allowed to vote in failover or accept reads until both phases have completed and its applied position is within a small bounded distance of the writer’s. There is no “promoted with stale state” path. This means a freshly-joined follower can take over from the writer the same millisecond it finishes catching up, without the hole that a naive file copy leaves behind. ## Why this is hard in practice The hard part isn’t the algorithm - every database textbook describes something like it. The hard part is making the snapshot phase consistent across many shapes (rows + vectors + secondary indexes + commit position + sequence values) without taking a global lock that stalls the writer. Our trick is the same substrate that handles atomic multi-shape writes: every write, no matter the shape, advances a single monotonic commit position. A “snapshot at position W₀” means: materialize the state as of W₀ and ship it. Because the materialization is a deterministic function of the committed history up to W₀, two followers that bootstrap from the same `W₀` produce bit-identical state. That last property - bit-identical materialization - is what lets followers agree on snapshot bytes during a write-heavy cluster join. ## The chaos drill Architecture is one thing; verified behavior is another. So we run the drill: 1. Spin up a 3-node cluster. 2. Start a writer that’s executing 10k writes/second across rows + vectors + indexes. 3. Snapshot-bootstrap a fresh follower mid-write. 4. Once the follower applies position == writer position, kill the writer. 5. Trigger failover. The follower is promoted. 6. Diff the promoted follower’s state against the committed writes the application reported success on. The pass condition is that the bootstrap leaves no hole: every write the follower was told about before it declared itself caught up must be present after promotion. The drill ran on 2026-04-30 against the production build and passed end-to-end. Note what the drill does not prove - it kills the writer only once the follower has caught up. Kill it while the follower is behind and you lose the difference, which is exactly the asynchronous-replication tradeoff described below. ## What this gives you - No seeding gap: the window a naive file-copy bootstrap leaves open - writes that land between snapshot start and snapshot end - is closed by construction. - Snapshot-bootstrap a follower against a hot writer without freezing the writer or pausing application traffic. - Self-healing replicas: a follower that fell too far behind is automatically re-snapshotted instead of replaying terabytes of change history. ## What it does not give you Being straight about the boundary: - This is not zero-data-loss failover. Replication to the standby is asynchronous: the primary streams committed changes continuously but does not wait for the standby before answering you. If the primary dies abruptly, the most recent acknowledged writes can be lost. How much is at risk depends on how far behind the standby had fallen. - What a 200 does guarantee is local durability: the write was flushed to durable storage on the primary before you got your response, so it survives a process crash or a host restart of that instance. - Cross-region replication. The writer and followers all live in one private network for now. Cross-region is on the roadmap after multi-writer lands. The practical advice that follows: make your writes idempotent, and retry through a failover. ## FAQ ### What is snapshot bootstrap? Snapshot bootstrap is the procedure a new database replica uses to catch up to the current writer’s state. In OriginChainDB it has two phases - copy a consistent snapshot at commit position `W₀`, then apply every write committed after `W₀` - and the replica is gated from voting in failover until both phases complete. ### How is this different from streaming replication? Streaming replication assumes the replica already has a consistent base state. Snapshot bootstrap is what gets you that base state without freezing the writer. The two work together: bootstrap once, stream forever after. ### Does the writer block during snapshot bootstrap? No. The snapshot is materialized from the committed history up to position `W₀`, while the writer continues taking writes at positions `W₀+1`, `W₀+2`, … in parallel. The follower replays the gap during the catch-up phase. ### How long does failover take? The cutover itself is fast. The slow part is deciding the writer is really gone: the old writer’s claim has to expire and a grace window has to pass before the standby is eligible, and both of those are configured rather than fixed. Automatic promotion is off by default, and when it is enabled it refuses to promote a standby that is not fully caught up. That is the standard tradeoff in any automatic-failover system, and we would rather be slow than promote over a live primary. ### What if all replicas are behind? Failover requires enough replicas to be within a bounded distance of the writer. If everyone’s far behind, failover refuses and pages a human. Better to alert than to silently promote a stale follower. ## What to read next - [One database, every query shape](https://originchaindb.com/blogs/one-database-every-query-shape) - the substrate that makes a consistent multi-shape snapshot possible. - [Multi-node configurations](https://originchaindb.com/docs/multi-node) - how node count and per-node standbys are chosen. - [Fuzzing a database](https://originchaindb.com/blogs/850000-random-api-probes-a-day) - the continuous canary that guards the API surface. --- # How a graph database works vs joins - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/how-graph-databases-work Published: 2026-07-14T09:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal graphcypherdatabaseengineering # How a graph database works vs joins A relational database re-derives relationships with joins at query time; a graph database stores them as structure and walks them. When each one wins. OriginChainDB Jul 14, 2026 About 6 min read Data models / Graph ## Relationships become a path you can follow. 1. Account Customer → holds → account 2. Order Customer → placed → order 3. Product Order → contains → product 4. Person Customer → referred → person A graph makes declared relationships explicit. Traversal follows those relationships; relational queries can describe the same connections through joins. TL;DR - In a relational database a relationship is a value in a join table: every hop is re-derived at query time by matching keys. In a graph database a relationship is stored structure: every hop is a pointer you follow. That one difference is why multi-hop questions - who’s connected to whom, through what, within how many steps - stay fast on a graph as data grows, and why they get combinatorially painful as SQL joins. Set-shaped work - aggregations, scans, reporting - still belongs in SQL. ## Connected data has a shape A social network, a payment flow, a referral chain, a supply chain, an agent’s memory of which tools touched which records - these aren’t tables that happen to reference each other. They’re networks: things (nodes) connected by typed, directional relationships (edges), both carrying properties. The graph data model stores exactly that: - Nodes - entities: `(patient)`, `(provider)`, `(account)`, `(device)`. - Edges - typed, directed relationships: `(account)-[:PAID]->(merchant)`, `(doctor)-[:REFERRED]->(specialist)`. Edges are first-class: they have their own properties (amount, timestamp, weight). Nothing exotic so far - you can model this relationally. The difference is what the database does with it. ## How relational handles relationships SQL represents a many-to-many relationship as a join table: `referrals(from_id, to_id, date)`. The relationship exists as rows, and every query that uses it must re-derive the connection by matching keys - that’s what a join is. One hop, fine. But connected questions are rarely one hop: “Which providers are within three referrals of this one?” Relationally, that’s the join table joined to itself three times. With a good index each hop is a per-key B-tree probe - the relationship is still re-derived at query time, at log cost per edge, and on wide frontiers the planner may fall back to whole-table hash joins. The sharper pain is structural: intermediate result sets fan out combinatorially before being filtered back down; “within one to five referrals” means generating SQL per depth or writing a recursive CTE the planner treats as a black box; and every added hop compounds all of it. ## What a graph database does differently A graph store makes adjacency a storage primitive. Each node knows its edges directly - conceptually, following an edge is a lookup keyed by the node, not a scan-and-match over a global table. The property this buys is the whole point: Traversal cost is proportional to the part of the graph you actually touch, not to the total size of the database. “Friends of friends of Alice” visits Alice, her ~50 edges, and their ~2,500 edges - the same work whether the database holds a thousand users or a hundred million. In the join world that same question keeps getting slower as tables grow, because every hop re-derives membership against ever-bigger structures. Traversals compose into the operations connected questions are made of: neighbors, shortest paths, variable-depth expansion (`1..5` hops), and pattern matching - “find this shape anywhere in the graph.” ## Cypher in sixty seconds Graph queries read like the sentence you’d say out loud. Cypher, the most widely adopted syntax, draws the pattern: ```cypher // Who did Dr. Rao refer patients to, directly or through one intermediary? MATCH (d:provider {name: "Dr. Rao"})-[:REFERRED*1..2]->(p:provider) RETURN DISTINCT p.name ``` ```cypher // Fraud shape: does money leaving this account cycle back to it within 4 hops? MATCH path = (a:account {id: $suspect})-[:PAID*2..4]->(a) RETURN path ``` The second query is the famous one. Finding cycles in SQL means self-joining a payments table up to four times and comparing endpoints - a query you write once, hate forever, and watch degrade as the table grows. In Cypher it’s one line, and the exploration stays bounded to the suspect’s own up-to-4-hop neighborhood - no global join materialization. ## The advantages, in one place 1. Multi-hop cost tracks the answer, not the dataset. You pay for the neighborhood you explore. This is the structural advantage; the rest follow from it. 2. Queries read like the question. Less translation between product thinking and query text means fewer wrong queries. 3. Variable and unknown depth are native. “Within N hops” and “shortest path” are expressions, not generated SQL. 4. Relationship-first evolution. New edge types are additive. The moment you need `(:device)-[:SHARED_BY]->(:account)` for fraud detection, you write edges - no join-table migration, no ORM churn. 5. Pattern detection as a first-class operation. Rings, chains, hubs, communities - shapes that signal fraud, influence, or risk - are match patterns, not analytics jobs. ## Where SQL still wins Graphs are not a better everything. Aggregations and scans - “revenue by region last quarter” - are set operations; SQL’s planner, indexes, and window functions are built for them and win. Constraint-heavy transactional flows lean relational. Nobody wants tax reporting as a traversal. The uncomfortable truth is that real applications need both shapes on the same data: the payment row that lands in your revenue report is the same payment edge the fraud query walks. Historically that meant running a relational database and a graph database, plus the ETL to keep two copies of the truth in sync - and that pipeline, not either database, becomes the thing that pages you. ## Graph on OriginChainDB Our answer is one store that speaks both: rows you query with SQL are the same records you traverse with Cypher - one write, no sync pipeline, no second system to operate. The referral queries above run against tables you can also `GROUP BY`; an agent’s tool-call history is rows for auditing and a graph for “what touched this record before it broke.” Vector similarity composes with both, because recommendation and memory questions are usually “nearest neighbors, then walk the graph from there.” If you’re weighing a dedicated graph database, the real question isn’t “graph or relational” - it’s whether the connected questions justify operating and synchronizing a second store. When the graph is another query shape over data you already have, the answer changes. ## FAQ ### Is a graph database just an ORM over join tables? No - the difference is physical, not syntactic. An ORM still emits joins that re-derive relationships per query; a graph store persists adjacency and walks it. ### How large does a graph get before traversals slow down? Traversal cost scales with edges visited, so the operative number is your data’s branching factor raised to the hop depth - exponential in depth, which is why bounding depth matters - not total graph size. Dense supernodes (a node with millions of edges) are the thing to model around, on any graph system. ### Do I have to learn a whole new query language? Cypher’s core - `MATCH`, patterns, `RETURN` - is learnable in an afternoon, and you keep SQL for everything set-shaped. The two coexist here deliberately. ### What are the canonical graph use cases? Fraud rings and collusion detection, recommendations, permission and dependency resolution, supply-chain tracing, knowledge graphs for agents, and agent memory - anywhere the question contains “connected to,” “through,” or “within N steps.” ## What to read next - [One database, every query shape](https://originchaindb.com/blogs/one-database-every-query-shape) - the architectural case for one substrate. - [Per-key TTL for agent memory](https://originchaindb.com/blogs/per-key-ttl-agent-memory) - agent memory patterns that pair naturally with traversal. - [How vector search actually works](https://originchaindb.com/blogs/how-vector-search-works) - the companion deep-dive on the vector query shape. --- # How vector search works: HNSW explained - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/how-vector-search-works Published: 2026-07-14T08:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal vector-searchembeddingshnswengineering # How vector search works: HNSW explained Embeddings turn meaning into geometry. A plain-language walk through similarity metrics, why brute force dies at scale, HNSW, and what quantization costs. OriginChainDB Jul 14, 2026 About 8 min read Data models / Vector search ## Search a neighborhood of similar vectors. 1. Embedding Represent the query numerically 2. Similarity Choose a distance metric 3. Index Explore promising neighbors 4. Results Return ranked record IDs Approximate nearest-neighbor search explores candidates in an index. Recall, query effort and memory are trade-offs to evaluate on your own data. TL;DR - An embedding model turns text (or images, or audio) into a point in a high-dimensional space, placed so that similar meanings land near each other. Vector search is “find the nearest points to this one.” Doing that exactly gets brutally expensive as your data grows, so production systems use approximate indexes - most famously HNSW - that answer in milliseconds while giving up a controlled, measurable amount of recall - recall@10 of 0.96 at p99 109 ms in our 100k-vector benchmark, or 0.69 at p99 37 ms on the same index with a narrower search frontier. ## From meaning to geometry An embedding model is a neural network with an unusual output: instead of a label or a sentence, it emits a fixed-length list of numbers - 384, 1024, 1536, sometimes 3072 of them. That list is a coordinate. Feed the model “How do I reset my password?” and “password recovery steps” and you get two coordinates that sit close together. Feed it “best pizza in Mumbai” and you get one far away from both. That’s the entire trick. The model was trained so that semantic similarity becomes spatial proximity. Once meaning is geometry, search stops being about matching words and starts being about measuring distance - which is why vector search finds “password recovery steps” for the query “reset my password” even though they share almost no tokens. Keyword search can’t do that; it was never told those phrases are the same idea. Every vector from a given model has the same dimensionality, and vectors from different models are not comparable - a 1536-dim OpenAI embedding and a 1024-dim Cohere embedding live in unrelated spaces. This is why the model you embed with is a schema-level decision, not an implementation detail (more on that at the end). ## Measuring “similar”: cosine, dot product, L2 Distance needs a definition. Three show up everywhere: - Cosine similarity - the angle between two vectors, ignoring their lengths. The default for text embeddings: you usually care about direction (meaning), not magnitude. - Dot product - angle and magnitude. Useful when the model encodes importance into vector length; also what you get for free if your vectors are normalized, since cosine and dot product are then identical. - Euclidean (L2) - straight-line distance. Common in vision and anywhere the space was trained with L2 losses. The practical rule: use whatever metric the embedding model was trained for. The model’s documentation says which; fighting it costs you recall for no benefit. ## Why brute force dies Exact search is trivial to write: compare the query against every vector, keep the top k. At small scale it’s genuinely fine: below a few tens of thousands of vectors, a straight scan is fast enough that an index isn’t buying you much - we keep an exact-scan path around for exactly that case, and for verifying index recall against ground truth. The problem is arithmetic. One million 1536-dim vectors means ~1.5 billion multiply-adds per query. Ten million means 15 billion. At useful traffic that’s not a tuning problem, it’s a physics problem - memory bandwidth alone puts a floor under your p99 that no amount of SIMD rescues. Exact search scales linearly with corpus size, and your corpus grows faster than your hardware budget. ## ANN: trading a little exactness for a lot of speed Approximate nearest neighbor (ANN) indexes accept a bounded error to escape the linear scan. The error is measured as recall@k: of the true k nearest neighbors, what fraction did the index return? Recall@10 of 0.96 means that on average 9.6 of the true top-10 made it into your answer. For semantic search this trade is almost always right. Embeddings are themselves approximations of meaning; the 4% of neighbors an index misses are overwhelmingly the marginal ones, not the obvious hits. What matters is that the trade is tunable and measured - which brings us to HNSW. ## HNSW: a highway system for vectors HNSW (Hierarchical Navigable Small World) is the index behind most production vector search today. Two ideas compose: Navigable small worlds. Connect each vector to a handful of its near neighbors and you get a graph you can traverse greedily: start anywhere, repeatedly hop to whichever neighbor is closest to the query, stop when no hop improves. Local links alone get stuck, so the construction also keeps some longer-range links - the “small world” property - letting greedy search cross the space in few hops. Hierarchy. Stack several such graphs. The top layer has few nodes and long links - think highways. Each layer down is denser and shorter-range - arterials, then streets. A query enters at the top, rides the highways to roughly the right neighborhood, then descends layer by layer, refining. The result is search cost that grows roughly logarithmically with corpus size instead of linearly. Two knobs matter in practice. M - how many links each node keeps - sets the memory/quality baseline of the graph. ef - how wide a frontier the search explores - is the live recall/latency dial: raise it for better recall, lower it for faster answers. In our benchmarks at 100k vectors, the default high-recall setting reaches recall@10 = 0.96 at p99 109 ms, and a fast mode answers at p99 37 ms where recall ≈ 0.69 is acceptable - same index, different ef. The point isn’t the specific numbers; it’s that the trade-off is explicit and yours to set. ([Benchmarks →](https://originchaindb.com/benchmarks)) ## Quantization: shrinking the vectors themselves The index solves “which vectors do I look at.” Quantization solves “how much does each look cost.” - float32 (none) - the raw embedding, 4 bytes per dimension. A 1536-dim vector is ~6 KB. - Scalar (int8) - each dimension squeezed to 1 byte: 4× less memory and bandwidth for a small, measurable recall cost - benchmark it on your own embeddings rather than trusting anyone’s blanket number. The workhorse. - Binary - one bit per dimension, 32× smaller. Distances become Hamming operations - extremely fast, at a real accuracy cost, so it’s best suited to candidate-generation stages that a finer pass then re-ranks. - Product quantization (PQ) - compresses whole sub-blocks of the vector via learned codebooks, 8-64× reductions. Paired with an inverted-file index (IVF-PQ) it’s how single machines serve nine-figure corpora: IVF-PQ compresses each vector so the index still fits one box. We have not published a measured result at that scale. The pattern across all four: spend memory where the search is fine-grained, compress where you’re only shortlisting. ## The part everyone forgets: filters Real queries are rarely “nearest 10 overall.” They’re “nearest 10 where tenant = X and status = active.” A filtered ANN query has to reconcile two machineries: the index proposes candidates by pure similarity, and the predicate disqualifies some of them. The standard mitigation - ours included - is candidate over-fetch: retrieve a deliberately larger set, apply the predicate, return the top k survivors. That holds recall up under moderately selective filters. Under needle-narrow filters - a predicate that passes one row in a million - no fixed multiplier saves you, and the honest plan inverts: filter first, then rank the survivors exactly. When you evaluate any vector system, ask how its filters behave as selectivity climbs; the answers differ more than the headline benchmarks do. ## What this looks like in OriginChainDB Vector search here isn’t a separate product to operate - it’s one query shape over the same store as your rows: - Per-table choice, made on first write: metric (cosine, dot, L2), index family (HNSW today; IVF-PQ for very large corpora), and quantization (none, int8, binary). Nothing to install, no ingestion job to babysit - vectors index on write. - One atomic write: the row and its embedding commit together, so a crashed pipeline can’t leave a document without its vector or a vector without its document. - Filters in the same query: metadata predicates are part of the topk call itself - enforced by the engine with candidate over-fetch, not left for you to bolt on after the results come back. - A model registry: because vectors from different models don’t mix, the console tracks which embedding model produced each table and drives [embedding migration](https://originchaindb.com/blogs/no-separate-vector-database) when you change models. ## FAQ ### Which metric should I pick? The one your embedding model documents. If vectors are normalized (most text models), cosine and dot product are equivalent - pick either and stay consistent. ### Do more dimensions mean better search? Higher-dim models often embed meaning more finely, but each dimension costs memory and bandwidth forever. Several modern models are trained so their vectors truncate gracefully; when yours is, benchmark the shorter length before paying for the full one. ### When is brute force the right answer? Small corpora - roughly under tens of thousands of vectors - and any case where exactness is contractual. A scan is exact and index-free, and at that scale an index isn’t buying you much. ### What happens when I switch embedding models? Every stored vector must be re-embedded - the old and new spaces are incomparable. Plan it like a schema migration: dual-write, backfill, cut over reads. ## What to read next - [Why we don’t need a separate vector database](https://originchaindb.com/blogs/no-separate-vector-database) - the architectural argument. - [RAG at 10M documents](https://originchaindb.com/blogs/rag-at-10m-documents-the-four-vendor-failure-mode) - where multi-system RAG stacks fail. - [The RAG latency budget](https://originchaindb.com/blogs/rag-latency-budget) - where those milliseconds actually go. --- # Idempotent tool calls for agent loops - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/idempotent-tool-calls Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal agentstool-callsidempotencyreliabilitytutorial # Idempotent tool calls for agent loops Agent tool calls with side effects get retried and duplicated. The idempotency-key playbook: atomic compare-and-set, per-key TTL, and wait-for-result. OriginChainDB Team May 5, 2026 About 6 min read Agent systems / Tool execution ## Track the operation before repeating its effect. 1. Identify Derive a stable operation key 2. Claim Coordinate one execution 3. Run Record the effect and result 4. Replay Handle a repeated call The application coordinates execution and result lookup around a stable identity. External side effects need their own idempotency strategy too. TL;DR - Production AI agents make tool calls that have side effects: sending emails, charging cards, posting webhooks. Network retries and LLM-generated duplicates mean you’ll get the same call twice, and the second one shouldn’t fire. The fix is an idempotency key derived from the intent, an atomic claim only one worker can win, and a stored result the losers read instead of re-firing. ## The problem Your agent decides to send an email. The HTTP call to your email service times out. The agent doesn’t know if the email was sent. It retries. Now the user gets two emails. Or: the LLM, in some non-determinism, hallucinates that it should call `send_email` twice in the same response. Your tool-handling code dutifully sends both. Or: a downstream queue retried the agent’s request after a 503. Same call, second invocation. Same problem. Every agent that touches the real world hits this. The fix is idempotency keys - every side-effecting call carries a unique key, and the system that processes it tracks which keys it’s already seen. ## The shape of an idempotent tool Three pieces: 1. The key. A stable identifier for the intent, not the call. If the agent decides “send email about order #12345 to user U”, the key should be a function of (intent type, order, user) - `send_email:order-12345:user-U`. Same intent, same key, regardless of how many times the call fires. 2. The store. A K/V store that records “I have seen this key, and the result was R.” Reads have to be fast (every call hits this); writes have to be atomic (no double-fires inside a race). 3. The compare-and-swap. When a call comes in, atomically: (a) check the store for the key, (b) if seen, return the stored result; (c) if not seen, claim the key, perform the side effect, store the result. Step (c) needs to be atomic. Otherwise two concurrent calls both see “not seen,” both claim the key, both fire the side effect. That’s the bug we’re trying to prevent. ## Building this on OriginChainDB Two OriginChainDB features make this clean: per-key TTL and atomic compare-and-swap. ```typescript async function handleToolCall(intent: ToolIntent): Promise { const key = idempotencyKey(intent); // e.g. "send_email:order-12345:user-U" // Check if we've already handled this intent const existing = await oc.get("tool-call", key); if (existing) { return existing.result; // already fired, return cached result } // Atomically claim the key with a short pending TTL // (so a crash before completion auto-clears the claim) try { await oc.put({ shape: "tool-call", id: key, set: { intent, status: "pending", started_at: Date.now }, ttl: 300, // 5 min pending window if_match: { exists: false }, }); } catch (e) { if (e.status === 412) { // Another worker claimed it first; wait for it to complete return await waitForResult(key); } throw e; } // We have the claim. Fire the side effect. let result: ToolResult; try { result = await fireSideEffect(intent); } catch (e) { // Mark failed; let the TTL clean up so retries can try again later await oc.put({ shape: "tool-call", id: key, set: { intent, status: "failed", error: String(e) }, ttl: 300, }); throw e; } // Persist the result. Long TTL so re-issued intents in the next // hour return the cached result instead of re-firing. await oc.put({ shape: "tool-call", id: key, set: { intent, status: "completed", result, completed_at: Date.now }, ttl: 3600, // 1 hour }); return result; } ``` The `if_match: { exists: false }` is the atomic claim. Because the substrate serializes the compare-and-swap, only one worker will succeed at the put; the rest get 412 and fall into the wait-for-result path. The TTLs do double duty: - The 5-minute pending TTL means if a worker crashes mid-call, the claim auto-clears and retries can try again. - The 1-hour completed TTL means re-issued intents within a reasonable window get the cached result; after that, the system has “forgotten” and will re-fire (which is usually fine - by an hour later the agent’s situation has changed). ## How to compute the key The key is a function of intent. Two design choices: Option A - derive deterministically from intent. `sha256(JSON.stringify(intent))`. Same intent always produces the same key. Good for “the LLM emitted the same call twice.” Option B - pass the key in from the caller. The agent’s loop knows it’s retrying and reuses the key. Good for “the upstream queue retried our handler.” In practice you want both. Compute a deterministic key from the intent, but let the caller override it. That way: - LLM duplicates (same intent, no caller key) deduplicate via the deterministic key. - Network retries (caller knows) reuse the explicit key. - Genuinely-new intents that look textually similar (rare) can still differ by the explicit key. ## The wait-for-result path When a second worker hits “I’d claim this but someone else already claimed it,” what should it do? ```typescript async function waitForResult(key: string, timeoutMs = 30_000): Promise { const start = Date.now; while (Date.now - start < timeoutMs) { const r = await oc.get("tool-call", key); if (r?.status === "completed") return r.result; if (r?.status === "failed") throw new Error(r.error); await sleep(100); } throw new Error("timed out waiting for in-flight tool call"); } ``` Polling is fine here - the second worker’s job is to wait, not to do work. 100ms intervals are cheap reads on a fast K/V store. For longer-running tools (LLM generation, etc.), bump the interval. For tools you know are slow (>5s), consider a real notification primitive - but for the typical email-send / webhook-post / db-update case, polling is correct and simple. ## What this gets you The agent’s tool layer becomes safe under: - LLM-driven duplicate calls (same intent, same key, dedupe via store). - Network retries (agent retries; reuses key; gets cached result). - Worker crashes mid-call (pending TTL clears; retry works). - Multiple agents converging on the same action (atomic claim ensures only one fires). What it doesn’t get you: protection against side effects that are themselves non-idempotent on the receiving end. If your email service charges per request regardless of dedup, you’re still paying. The discipline is “idempotency at every layer that has state” - the tool layer is one of them. ## FAQ ### What’s an idempotency key? A unique identifier for an intended operation, used to detect duplicate invocations. Same key = same intent = the system should fire the side effect once and return the same result on subsequent calls. ### Should I derive the key or pass it in? Both. Derive deterministically by default (handles LLM duplicates), but accept a caller override (handles retry scenarios where the caller knows it’s retrying). ### What’s the right TTL for completed tool calls? Long enough to cover all reasonable retry windows from the upstream caller, short enough that storage doesn’t grow forever. 1 hour is a reasonable default; tune based on your retry policy. ### Does this work for streaming tool calls? Streaming makes idempotency harder - partial results don’t fit the “completed once” model cleanly. The pragmatic answer: serve the stream from the original handler, and on retry, return the final result (cached) rather than re-streaming. ### Does OriginChainDB’s `if_match` add latency? ~20-50µs over a non-CAS write. Negligible compared to the side effect itself. ## What to read next - [Automatic Idempotency-Key in the SDK](https://originchaindb.com/blogs/auto-idempotency-key-sdk) - the same guarantee, done for you on every HTTP write. - [Per-key TTL](https://originchaindb.com/blogs/per-key-ttl-agent-memory) - the TTL semantics this design relies on. - [Write backpressure: 429 + Retry-After](https://originchaindb.com/blogs/backpressure-429) - the retries this design has to absorb. --- # One MCP server for every AI IDE - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/mcp-server-launch Published: 2026-05-04T10:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal mcpai-toolsclaudecursor # One MCP server for every AI IDE The @originchain/mcp-server package exposes five tools over stdio transport, configured by environment variables. What it does, and what MCP doesn't solve. OriginChainDB May 4, 2026 About 6 min read Integrations / MCP ## Connect a coding tool to database context. 1. AI client Choose a database tool 2. MCP server Expose the tool interface 3. Database API Execute an authorized request 4. Tool result Return context to the client MCP carries tool requests and results between a client and server. Credentials, permissions and application decisions still define what the tool may do. The OriginChainDB MCP server exposes five tools - oc_ask, oc_sql, oc_vector_topk, oc_fts_search and oc_list_schemas - so an agent inside an AI IDE queries your tenant directly. It runs over stdio as a child process of the IDE, reads three environment variables, and keeps no state of its own. If you watched the AI tooling stack settle in 2025–2026, the tooling layer that won is MCP - the Model Context Protocol. It is a thin JSON-RPC contract that lets a client (Claude Desktop, Cursor, Zed, Windsurf, Continue, OpenAI’s Apps SDK) plug into a server (a database, a filesystem, a ticket tracker) over stdio or HTTP. The agent inside the IDE sees the server’s tools - typed, named functions with JSON-Schema arguments - and calls them when the user’s prompt suggests it should. This is not a niche developer-experience improvement. It is the new front door. When an engineer opens Claude Desktop and asks “what’s the schema of the products table?”, the answer comes from whichever database server is registered in `claude_desktop_config.json` - and the database that isn’t there at all does not get queried, recommended, or even mentioned. Database vendors without an MCP server are, increasingly, vendors who don’t exist inside the agent’s tool list. We shipped `@originchain/mcp-server` for exactly this reason. ## The five tools An MCP server’s surface is its tool list. Tools are named functions; the agent reads the name + description + JSON-Schema and decides when to call. Bigger surfaces are not better - every tool the agent considers is an opportunity to call the wrong one. We exposed five, mapping 1:1 to the OriginChainDB query shapes: ```plaintext oc_ask Natural-language question → rows back oc_sql Typed SQL with bind params oc_vector_topk HNSW similarity search with metadata filter oc_fts_search BM25 full-text search oc_list_schemas Introspect the tenant's tables, columns, indexes ``` `oc_list_schemas` matters more than it looks. Without it, the agent has to guess column names and table layouts; with it, the very first call from a fresh session produces a schema the agent can reason from. We see this in traces - Claude consistently calls `oc_list_schemas` first when the user’s question references a table the agent hasn’t seen, then writes a `oc_sql` query that’s right on the first try. Each tool’s input schema is small and typed. `oc_sql` takes `sql` (string) and `params` (array). `oc_vector_topk` takes `table`, `vector`, `k`, optional `mode` (`fast` or `high_recall`), and optional `filter`. The schemas are derived from the same OpenAPI spec the SDK is built from, so the agent sees the same shapes a TypeScript client would. ## The stdio transport pattern MCP supports two transports - stdio and HTTP. For local IDE integrations, stdio is the right answer. The IDE spawns the server as a child process; JSON-RPC messages travel over stdin/stdout; lifecycle is bound to the IDE’s lifecycle. No port to allocate, no localhost server to forget about, no auth surface beyond what the IDE itself enforces. The server’s main loop is small: ```ts import { Server } from "@modelcontextprotocol/sdk/server/index.js"; import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js"; const server = new Server( { name: "originchain", version: "0.1.0" }, { capabilities: { tools: {} } } ); server.setRequestHandler(ListToolsRequestSchema, async => ({ tools: [OC_ASK, OC_SQL, OC_VECTOR_TOPK, OC_FTS_SEARCH, OC_LIST_SCHEMAS], })); server.setRequestHandler(CallToolRequestSchema, async (req) => { switch (req.params.name) { case "oc_sql": return runSql(req.params.arguments); case "oc_ask": return runAsk(req.params.arguments); //... } }); await server.connect(new StdioServerTransport); ``` Every handler hits the tenant’s HTTPS endpoint with the bearer token from the env. There is no local cache, no connection pool, no state - the MCP server is a thin shim, and that’s the point. When the engine ships a new query shape, the spec updates and the shim rebuilds; there is nothing in the shim to break. ## Env-var config The server reads three environment variables: ```plaintext ORIGINCHAIN_URL https://acme.originchain.ai ORIGINCHAIN_TOKEN bearer_xxxxxxxxxxxxxxxxxxxxxxxx ORIGINCHAIN_TENANT acme # optional; inferred from URL when absent ``` No config files. No interactive setup. The IDE’s config block holds the env, the server reads it on launch, and there is exactly one place to rotate the token when it expires. This matters for security review - tokens never land on disk in a config file the user might commit. ## A working Claude Desktop config Drop this into `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS) or the platform equivalent on Windows / Linux: ```json { "mcpServers": { "originchain": { "command": "npx", "args": ["-y", "@originchain/mcp-server"], "env": { "ORIGINCHAIN_URL": "https://acme.originchain.ai", "ORIGINCHAIN_TOKEN": "bearer_xxxxxxxxxxxxxxxxxxxxxxxx" } } } } ``` Restart Claude Desktop. Open a fresh chat. Ask “what tables do I have?” and the agent calls `oc_list_schemas`. Ask “top 10 suppliers in Mumbai by shipment volume last quarter” and it calls `oc_ask`, which routes through the natural-language planner on the tenant. Ask “find products similar to this paragraph and group by category” and it calls `oc_vector_topk` followed by `oc_sql` to aggregate. The whole loop closes inside the IDE. Cursor’s setup is identical - `~/.cursor/mcp.json`, same shape, same env. We tested against both clients on the same tenant. ## What MCP doesn’t yet solve Three things the protocol is honest about not solving, and that matter for production database integrations: - Agent observability. When the agent calls `oc_sql` ten times in a session, who sees those calls? The IDE logs them locally; the database logs them under the bearer’s tenant; there is no per-agent identity. If your auditor asks “which agent ran which query?” the answer today is “the bearer”, not “the agent.” Per-agent identity over MCP is being discussed; it is not in the wire format yet. - Granular permissions. `oc_sql` either runs or doesn’t. There is no row-level scope inside the tool call, no “read-only on these tables” gate, no rate limit per tool per session. The bearer’s tenant-scope is the only fence. We mitigate this by recommending a separate read-only bearer for IDE use, but the protocol itself doesn’t have a permission story. - Billing-per-call. The free database caps `/ask` at 50 calls per day. MCP has no notion of cost. An agent that loops `oc_ask` 200 times in a runaway session burns that quota with no IDE-side warning. Quota is enforced server-side via 429 responses; the agent will see them and back off, but the cost-control story is one-way (the server tells the client to stop) rather than budgeted (the client knows in advance). None of these are blockers for shipping. They are the shape of the next protocol revision, and the kind of work the database vendor gets to do when the protocol catches up. ## Status The MCP server is built and tested against Claude Desktop and Cursor on macOS. It is not yet on npm publicly - the GitHub repo at `github.com/originchain-ai/originchain-mcp` is being pushed in the next release window. Until then, paid tenants who want early access can email support and we’ll wire them up directly. Provision a tenant in under two minutes via the [quickstart](https://originchaindb.com/docs/quickstart). The MCP repo lands at `github.com/originchain-ai/originchain-mcp` soon - bookmark it. One bearer. One server. Every IDE. --- # Why the database is now multi-modal - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/multi-modal-is-the-new-foundation Published: 2026-07-01T08:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-14T15:11:01.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal multi-modalaivector-searchagentic-aienterprise # Why the database is now multi-modal SQL, vector, graph, full-text and natural-language queries usually mean five systems kept in sync. A multi-modal database answers all five on one copy. Zaheer Jul 1, 2026 About 3 min read Architecture / Multimodal data ## Different questions need different representations. 1. SQL Filter and aggregate fields 2. Vector Find semantic similarity 3. Graph Follow declared relations 4. Full-text Match words and phrases Four data models serve different questions. ASK is a query interface; your application supplies embeddings and coordinates the required index writes. A multi-modal database answers SQL, vector, graph, full-text and natural-language queries over one copy of the data, on one endpoint. That matters because AI workloads need all five shapes at once, and every seam between five separate systems is a place where the copies drift apart. For decades, the database had one job: store structured data and return exact answers to precise queries. That model served the transactional era well. It is no longer enough for the AI era. ## 1. Agents are moving to the center Within the next three to five years, people will increasingly stop operating data platforms directly. Instead, they’ll orchestrate AI agents that query, reason over, and act on data on their behalf. Increasingly, the database’s primary “user” is a machine. ## 2. Queries are becoming flexible The rigid, exact-match query is giving way to natural-language, context-aware retrieval — questions that carry the memory of previous interactions and aim for the best answer, not merely a literal one. ## 3. Infrastructure is going AI-native Real intelligence needs more than one data shape. Vector embeddings for semantic meaning. Full-text for language. Graph for relationships. Structured tables for facts. AI workloads need all of them — together. ## This is exactly where most of today’s stacks break To ship AI experiences, enterprises are currently forced to bolt together separate systems — a relational database here, a vector store there, a graph engine, a search cluster, and an orchestration layer to keep them in sync. Every seam adds latency, cost, and operational overhead — and, most dangerously, another place where data drifts out of consistency. And there’s a deeper contradiction this fragmented world cannot resolve: Flexible queries are the whole point — but some answers must be exact. A fraud decision. A trade settlement. A patient record. A regulatory filing. When a workload carries legal or financial weight, a “best guess” is a liability. The AI-era database has to deliver semantic flexibility and deterministic precision — from the same system, on the same data. ## This is why we built OriginChainDB as a true multi-modal database Not a relational engine with a vector add-on. Not a search index pretending to be a database. A single, AI-native platform — engineered in Rust — where SQL, Vector, Graph, Full-Text, and Natural Language all run on one endpoint: - Semantic and natural-language retrieval when you need flexibility - Deterministic, exact SQL when correctness is non-negotiable - Sub-100 ms latency at 10,000+ QPS, backed by an availability SLA of up to 99.95% on high-availability configurations - Single-tenant dedicated configurations in the region you choose, or on-prem — your data never leaves your control One endpoint. One consistent copy of your data. Flexibility where you want it, exactness where you need it. The agentic future doesn’t need five databases stitched together. It needs one that was designed for it. --- # Multi-region cold standby at $5/month - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/multi-region-cold-standby-5-vs-78-per-month Published: 2026-06-07T09:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal drinfrastructurecostengineeringoperations # Multi-region cold standby at $5/month How we cut per-tenant multi-region DR from ~$78/month to ~$5–15/month by dropping the standby host, and what it costs: RTO ~10–15 min, RPO ≤60 s. OriginChainDB engineering Jun 7, 2026 About 9 min read Operations / Recovery design ## A recovery path includes time to rebuild. 1. Primary Run the active workload 2. Archive Retain recovery artifacts 3. Provision Bring up recovery capacity 4. Restore Recover, validate and redirect Cold standby trades continuously running capacity for provisioning and restoration work. Historical costs and recovery estimates are specific to the case study. TL;DR — We redesigned our multi-region disaster recovery from a textbook hot-standby model to a cold-standby model and dropped per-tenant DR cost from ~$78/month to ~$5–15/month. The trade-off is ~10–15 minutes of RTO instead of ~5. For a single-customer pilot running on a managed cloud, that’s the right call. Here’s the math, the runbook excerpt, and the things this design can’t do. If you’ve read any production-database playbook from the last decade, the hot-standby DR model is the default answer: in every region you might ever fail over to, keep a continuously-running follower replica. It tails the writer’s log in near-real-time. When the primary region goes down, you promote the follower, flip DNS, and the customer sees a 5-minute outage instead of a 5-hour one. It is the right answer when you have many customers, real revenue per minute, and a contractual RTO measured in minutes. It is the wrong answer when you have one pilot customer, the engine has been live for six weeks, and the standby host costs more per month than the active customer pays you. ## The hot-standby model we started with The first version of multi-region DR, shipped a few weeks ago, ran the textbook pattern: - Per enrolled tenant, one dedicated host in the DR region, running the engine binary as a follower replica. - The follower replicated the writer’s committed changes via cross-region encrypted object storage (replicated snapshots and incremental state). - A failover health check, with a 90-second detection window, watched the primary’s `/health` endpoint and flipped a DNS failover record when the primary went silent. - On promotion, the operator promoted the follower to writer and traffic resumed inside 5 minutes. This worked. It passed the chaos drill. It was the design we’d recommend to a Series A customer with five-nines aspirations. It cost about $78/month per enrolled tenant, per region of failover coverage. Most of that was the follower host itself: a steady-state mid-size compute instance, sitting in the DR region, running the binary, consuming bytes from the replicated log, and serving exactly zero customer traffic. Across a single canary tenant, that’s ~$78/month. Across ten tenants, $780/month. Across one hundred, $7,800/month — a real line item on the cloud bill, for a feature that fires once a year, maybe. ## Why this is wrong for a single-customer pilot The right way to think about DR cost is “cost per outage-second avoided.” The hot-standby model buys you ~5 minutes of RTO. The cold-standby model — the one we shipped this week — buys ~10–15 minutes. The delta is 5–10 extra minutes of customer impact per regional outage. For a single pilot customer who hasn’t paid us a dollar yet, on an infrastructure that’s seen one (zero) regional outage in its lifetime, paying $78/month every month for the 5-minute RTO instead of the 15-minute RTO is the wrong shape of bet. We’re insuring against an event that hasn’t happened, for a customer who hasn’t paid us, at a steady-state cost that compounds with every tenant we onboard. The bet that is the right shape: keep all the cheap scaffolding (pre-allocated routable IP, DNS failover record, role, firewall rules, alarms, the replicated-change storage) so a future failover is a one-line operator command. Drop the standby host. Conjure it at failover time from the pre-baked machine image we already keep in the DR region. That’s the cold-standby refactor. ## What we kept hot, what we made cold The bill of materials, after the refactor: | Component | State | Why | | --- | --- | --- | | Cross-region change replication | Hot | This is the RPO. Drop it and we lose data on failover. | | Pre-allocated routable IP in the DR region | Hot | DNS points at this IP throughout. Don’t want to reallocate at failover time — DNS would have to wait for propagation twice. | | DNS failover record | Hot | Pre-wired with the failover IP as the secondary answer. Promotion just trips the health check. | | Pre-baked machine image | Hot | Already in the DR region, tagged identically to the primary. Conjures the host in ~3 minutes. | | Role, firewall rules, networking | Hot | All pre-allocated. Trivially cheap to keep around. | | Health check, alarm, paging topic | Hot | Same as before. The DR-region health check is what trips DNS. | | The standby host itself | Cold | This is the line item we cut. ~$58–115/month per tenant per region. | The economics: - Steady-state cost per enrolled tenant: ~$5–15/month for the scaffolding (routable IP, replication transit, alarm, health check, certificate storage). - Failover-only cost: $78/month-equivalent (prorated to the actual hours the conjured host runs). - Time to first byte served from DR on failover: 2–3 minutes of host launch + 60–90 seconds of bootstrap + 30 seconds for the engine to bind + DNS propagation. End-to-end RTO ~10–15 minutes with normal on-call response time. For a single canary tenant, we’re now paying ~$5/month for DR coverage instead of ~$78/month. That’s the headline. ## The operator runbook, excerpted The failover procedure, after the refactor, has one new load-bearing step at the front: spin up the host. The shape of it: ```plaintext # Add the tenant to the active-DR list. If multiple tenants are # already in failover, APPEND — don't replace. activate-dr --tenant --active ``` The active-DR list is the switch. In steady state it’s empty: the DR-side host is provisioned per entry in that list, which evaluates to zero hosts and costs zero dollars. At failover time, the operator appends the tenant id, applies the change, and the host materializes from the pre-baked image, attaches to the pre-allocated IP, and starts up as a follower replica against the replicated change stream. About 2–3 minutes for the host to launch, 60–90 seconds for the bootstrap script, and 30 seconds for the engine to bind to its HTTPS listener. The DNS failover record has already been pointing at the routable IP throughout — the DNS layer doesn’t need to change anything at promotion time. As soon as the engine answers `/health`, traffic flows. Promotion is the second step: promote the follower to writer. That’s another ~30 seconds. End-to-end RTO from the moment the primary region goes silent: ~10–15 minutes with normal on-call response, ~5 minutes if the operator was already at the terminal. ## RPO is unchanged The cold-standby refactor doesn’t change RPO. Our recovery point is still bounded by the lag of the cross-region change replication, and it assumes off-box change shipping is enabled for the tenant - it is opt-in and in preview, and without it the recovery point falls back to the daily storage snapshot — typically ≤60 seconds during normal traffic, with outliers up to a few minutes during multi-region congestion in the underlying infrastructure. A customer on a high-write-rate workload at the moment of a regional outage should still expect MB-scale data loss within the RPO window. That’s a property of replicated object-storage-backed change shipping, not of whether the standby host is hot or cold. Dropping the host doesn’t make replication slower. ## What the refactor can’t do Three things, in order of importance: - Sub-5-minute RTO is off the table. The host needs to boot. That’s 3 minutes minimum, on the best day, with everything pre-warmed. If your contractual SLA is “minutes” measured in the low single digits, cold standby is the wrong design and we’d need to go back to hot. - Concurrent failover of many tenants takes longer. If ten tenants all need DR hosts at once, the apply runs ten host launches in parallel — but per-account provisioning rate limits mean the actual provisioning isn’t ten-at-once. We’ve sized the runbook for one tenant at a time. For an event where many tenants fail over together, the procedure would need to batch. - No protection against a DR-region failure-domain outage at the same moment. v1 of the DR design already had this gap: the DR host lives in a single failure domain of the DR region. If that domain outages at the same instant as the primary region, the DR is down too. Probability is roughly 10^-7 per hour; we haven’t budgeted to fix it. The hot-vs-cold refactor doesn’t change this property. The first one is the load-bearing trade. We named it explicitly in the runbook header, so a 3am on-call operator doesn’t get surprised: “the standby host is conjured at failover time. Saves ~$78/month per enrolled tenant per region. RTO is ~10–15 min instead of ~5 min. The new load-bearing first step at failover time is to spin up the host.” ## When to flip back to hot The refactor isn’t permanent. The switch to flip back to hot standby is one line: set the active-DR list to the full enrollment list, and the host gets pre-allocated again, paying the $78/month every month from then on. The criteria we’d use to flip a specific tenant from cold to hot: - Contractual RTO ≤ 5 minutes. - Sustained write volume where the RPO outlier scenario (several minutes of replication lag during a busy minute) would cost more than $78/month in customer trust. - A history of regional outages in the primary region — empirical not theoretical. Until any of those is true for a given tenant, cold is the right shape. The DR design supports both modes; the decision is per-tenant, not global. ## The math, one more time The cost-per-outage-second-avoided math, written out: - Hot standby: ~$78/month × 12 months = $936/year, for ~5-minute RTO. - Cold standby: ~$5/month × 12 months = $60/year, for ~15-minute RTO. If a regional outage happens once per year: - Hot standby costs $936 to avoid ~10 minutes of customer impact = $93.60/minute avoided. - Cold standby costs $60 to take the full ~15 minutes = $4/minute “saved” by not paying. For a paying enterprise customer where $93.60/minute is cheap insurance, hot is right. For a single pilot canary, $4/minute saved is the right shape of bet — especially given how few regional outages there actually are in a year. This is the discipline of pricing infrastructure against the customer paying for it, not against the customer you wish you had. Today’s tenants get cold. Tomorrow’s might get hot. The DR design supports both with a one-line change. ## Try it If you’re building on OriginChainDB, this DR posture is what your tenant inherits today: cold standby in a second region, cross-region change replication on, DNS failover wired, 10–15 minute RTO documented. Create a [free database](https://originchaindb.com/docs/quickstart) and the same posture applies to your tenant from day one. When you graduate to a tier that needs hot standby, the same module flips on with one variable. --- # Do you need a separate vector database? - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/no-separate-vector-database Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal architecturevector-databasedesignai-native # Do you need a separate vector database? A separate vector database means dual-writing the entity and its embedding. Storing vectors as another shape in one substrate makes both writes atomic. OriginChainDB Team May 5, 2026 About 5 min read Architecture / System choice ## Evaluate the full retrieval path. 1. Source database Own the original record 2. Vector service Maintain retrieval entries 3. Multimodal engine Keep supported models together 4. Application Coordinate ingestion and results Compare ownership, query requirements and operations. Sharing a database does not mean an ordinary row write automatically submits its embedding. TL;DR - A separate vector database is the wrong shape for AI-native applications. Vectors describe entities; entities have schemas, lifecycles, and consistency needs. Splitting them across two systems creates dual-write bugs, replication lag, and operational overhead with nothing to show for it. We argue the correct architecture is one substrate, multiple shapes - and that this is the entire premise of OriginChainDB. ## What a “vector database” actually is A vector database stores high-dimensional vectors and supports approximate nearest-neighbor (ANN) search over them. Most also support filtering on metadata attached to each vector. Pinecone, Weaviate, Milvus, Chroma, Qdrant - they’re all variations on this theme. The thing they have in common: they assume the rest of your data lives somewhere else. The vector is a sidecar to a “real” record that’s stored in Postgres, Mongo, DynamoDB, or wherever your application’s source of truth lives. You generate the embedding from the real record, push it to the vector DB, and rely on a stream or worker to keep the two in sync. This is the architecture we think is wrong. ## What it costs you Dual-write bugs. Every team running this architecture eventually debugs a “why isn’t this record in search results?” report. The entity was written to the primary DB, but the embedding generation failed (or was retried, or the vector DB rate-limited). Some window of writes are visible in the primary but invisible to ANN search. You can mitigate this with idempotent workers, dead-letter queues, and reconciliation jobs - but you can’t eliminate it without atomic writes across both systems, which the architecture rules out by definition. Operational duplication. Two systems means two backup strategies, two failover plans, two monitoring dashboards, two on-call rotations (or one rotation responsible for two), two cost models. Each system has its own scaling story you have to learn. Schema drift. The vector DB has its own metadata schema. Adding a field to your entity that you want to filter ANN results on is two operations: update the entity schema, then update the vector DB metadata schema. They must be kept consistent forever. They drift. Cost duplication. You’re paying to store the same identifying information twice - once on the entity in the primary DB, once as metadata on the vector. At small scale this is invisible; at hundreds of millions of records it’s not. ## What the alternative looks like We model vectors as just another shape in the substrate. ```yaml - shape: article key: "article/{id}" value: { title: string, body: string, author: string } - shape: article-embedding key: "vec/article/{id}/body" value: { type: f32, dim: 768 } ``` Two shapes. One substrate. One atomic write. When you write an article, you write its embedding in the same transaction: ```typescript await oc.transaction([ { shape: "article", id, set: { title, body, author } }, { shape: "article-embedding", id, set: embedding }, ]); ``` Both writes commit in a single atomic operation. The article is never visible to a regular read without the vector being visible to ANN search, and vice versa. There is no replication lag because there is no replica. When you ANN-search, you get back ids; you hydrate them with a parallel batch read against the entity shape. One round-trip from the client, two reads inside the substrate. ## ”But isn’t a specialty engine faster?” This is the standard objection. Specialty vector databases tune their ANN engines harder than a general-purpose substrate would. Surely they’re faster. The answer is: yes, at very large scale (>100M vectors with high QPS), and no, anywhere else. We measure ANN search at sub-millisecond p50 on corpuses up to ~10M vectors. Past that, the gap to a dedicated engine starts to widen, and the depth-first optimiser work on our roadmap is what closes it. The teams running into the “vector DB is faster” wall are usually past 100M vectors. Most teams aren’t there. And even at that scale, the cost of running two systems is rarely paid for by the marginal latency improvement on the search call. ## ”But what about graph queries / full-text / time-series?” Same argument. Each of these is a “specialty engine” in the same sense a vector DB is. The right architecture isn’t to bolt them on as separate services - it’s to model them as shapes in one substrate. We’re not there for graph or FTS yet (those are post-1.0; see [the roadmap](https://originchaindb.com/blogs/depth-first-roadmap-to-1-0)). The principle stands: one substrate, multiple shapes. ## When the separate-vector-DB model still wins Honest answer: when your primary database is non-negotiable and your vectors are an add-on feature you’re prototyping. If you have a 10-year-old Postgres deployment, established team workflows, and a small AI feature being trialed on the side, adding Pinecone is the lowest-disruption path. The architecture is wrong, but the migration cost would be wronger. If you’re greenfield, or your AI feature is core to the product, or you’re already feeling the dual-write pain - that’s when one substrate becomes the right answer. ## FAQ ### What is a vector database? A vector database stores high-dimensional vectors (typically embeddings from ML models) and supports approximate nearest-neighbor search over them. Most also support metadata filtering. Examples: Pinecone, Weaviate, Milvus, Qdrant, Chroma. ### Why is a separate vector database the wrong architecture for AI-native apps? Because it requires dual-writing the entity (in your primary DB) and its embedding (in the vector DB). This creates eventual-consistency windows, operational duplication, and schema drift between the two systems. AI-native applications produce so many writes per session that these costs compound quickly. ### Doesn’t OriginChainDB just contain a vector engine? Yes - but the engine reads from the same substrate as every other shape. The substrate is unified; the vector engine is a query path, not a separate system. ### Is the unified-substrate model novel? The principle isn’t novel - Postgres + pgvector is the same idea. What’s novel is doing it from a managed K/V substrate purpose-built for AI workloads, with the throughput and atomic-multi-shape characteristics that follow. ### Can I keep using a separate vector DB if I want? Of course. OriginChainDB isn’t trying to take over your stack - it’s an alternative to the dual-write architecture. If your team is happy with the separate-vector-DB model, stay there. ## What to read next - [OriginChainDB vs Pinecone](https://originchaindb.com/blogs/vs-pinecone) - the head-to-head comparison. - [Schemas](https://originchaindb.com/docs/schemas) - the data model that makes the unified substrate work. - [Transactions](https://originchaindb.com/docs/transactions) - how a multi-shape write commits atomically. --- # One database, every query shape - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/one-database-every-query-shape Published: 2026-06-22T08:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-14T18:28:08.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal databaseaivector-searchengineering # One database, every query shape OriginChainDB answers SQL, vector, full-text, graph and natural-language queries from one managed store, with the same write visible to every shape atomically. OriginChainDB Jun 22, 2026 About 4 min read Architecture / Query models ## One engine, several ways to ask. 1. Records SQL over registered schemas 2. Meaning Vector similarity search 3. Relationships Graph traversal 4. Language Full-text word matching The native query models are SQL, vector, graph and full-text. Natural-language ASK provides another query interface, not a fifth storage model. OriginChainDB is one managed database that answers five query shapes against the same rows: SQL, vector similarity, full-text, graph traversal, and natural language. The same write is visible to every shape atomically, so there is no sync job standing between a database and a vector index. The alternative is what most AI applications end up running. You start with a relational database for your rows. Then you need similarity search, so you bolt on a vector database. Then keyword search, so a full-text index. Then a recommendation feature wants graph traversal, so that’s a fourth system — and most of your engineering time goes into pipelines that copy data between four stores and pray they stay in sync. The five shapes: - SQL — typed `SELECT`s, joins, aggregates, window functions. - Vector search — top-k similarity with cosine, dot, and L2. - Full-text — BM25 ranking with stemming across 18 languages. - Graph — BFS, shortest-path (Dijkstra), reverse traversal, PageRank. - Natural language — ask a question in plain English; the engine plans and runs it. One bearer token, one URL, one bill. No second store to sync. ## The part that actually matters: atomicity The reason teams tolerate the multi-system stack is that nobody has made the single-system version consistent. The whole point of OriginChainDB is that the same write is visible to every query shape, atomically. Insert a product row and its embedding in one request, and a vector search sees it the instant a SQL query does. There’s no replication lag between your “database” and your “vector index,” because there is no second copy — the vector lives next to the row it describes. That’s the difference between “we support vectors” and “vectors are a first-class shape of the same store.” ## A concrete example Say you’re building product search. A user types “lightweight running shoes under $120.” On a typical stack that’s three systems and a coordinator. On OriginChainDB it’s one store: - BM25 full-text over the product descriptions for the keywords, - a vector top-k for semantic similarity to the query embedding, - a SQL predicate for `price < 12000`, all against the same rows, returning JSON. When you add a product, every one of those shapes sees it immediately — no nightly reindex, no sync job to monitor. ## On performance, with receipts We don’t think you should take “it’s fast” on faith, so the runs we have published are on [/benchmarks](https://originchaindb.com/benchmarks). Vector search runs HNSW for low-latency recall and IVF-PQ for large corpora, where compressing each vector keeps the index on a single box. We have not published a measured result at 100M scale. For smaller corpora where latency is everything, the HNSW path runs at single-digit-to-low-double-digit-millisecond p99. In-region typed queries return under ~100 ms p99. ## Natural language that compiles to a plan, not a prompt The `/ask` endpoint is not a chatbot wrapper around your data. A plain-English question compiles to a real query plan — the same plan tree SQL and vector queries use — which you can inspect before it runs. Ask the same shape twice and the plan is cached, so the repeat skips compilation and returns warm in tens of milliseconds. It’s natural language as a front-end to the planner, not a prompt tax bolted on top. ## What’s honest to say we don’t do yet We’d rather name the gaps than have you find them: - Managed embeddings are still bring-your-own-vectors today. You compute embeddings on your side and write the vectors; we store, index, and search them. Server-side text→vector is on the roadmap. - SQL is a growing subset. Joins, aggregates, window functions, and subqueries work; recursive CTEs and `UNION` don’t yet. Every unsupported shape fails with a clear message that tells you the workaround, not a silent wrong answer. ## Try it OriginChainDB is a managed, region-isolated database. Start free with SQL, vector, full-text and graph, or provision a dedicated configuration to add natural-language `/ask`, and try them from the [quickstart](https://originchaindb.com/docs/quickstart), or read how the engine answers each shape in the [architecture guide](https://originchaindb.com/architecture). OriginChainDB is built by Silicoyn Technologies. Learn more at [originchaindb.com](https://originchaindb.com/). --- # OpenAPI spec and the AI coding loop - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/openapi-spec-and-the-ai-coding-loop Published: 2026-05-04T13:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T20:18:50.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal openapiai-toolssdkdeveloper-experience # OpenAPI spec and the AI coding loop An OpenAPI spec is what a coding agent reads to generate a client for your API. Ours is vanilla OpenAPI 3.1, one bearer scheme, published at /openapi.json. OriginChainDB May 4, 2026 About 6 min read Developer tools / OpenAPI ## Make the API contract part of the coding loop. 1. Read Load the published specification 2. Generate Build a typed request 3. Validate Check the response and errors 4. Refine Keep code aligned with the API An API specification describes request and response shapes. Runtime limits, deployment state and application behavior need separate verification. OriginChainDB publishes a vanilla OpenAPI 3.1 spec at `/openapi.json` because that is the file a coding agent reads when an engineer asks it to build a client. Nobody filed an issue asking for one. The spec is the floor the whole SDK-generation loop stands on, and a vendor without one is not reachable from inside the agents engineers actually use. ## The SDK-generation loop Here is the workflow, and it is roughly universal now: 1. An engineer wants to call OriginChainDB from a Go service. 2. They open their coding agent and type: “build me a Go client for OriginChain. Here’s the OpenAPI spec: [https://acme.originchain.ai/openapi.json](https://acme.originchain.ai/openapi.json).” 3. The agent fetches the spec, reads the path patterns, response shapes, auth scheme, and error envelope, and either: - runs `openapi-generator-cli generate -i... -g go -o./oc-client`, then patches the bits the generator gets wrong, or - generates the client from scratch with full control over package layout, naming, and idiom. 4. Thirty seconds later, the engineer has a working `oc.NewClient(token).SQL(ctx,...)` they can drop into their service. This loop runs hundreds of times a day across every database vendor that has a spec. The vendor never sees it - there is no telemetry, no sign-up, no email - but it is happening, and it is the dominant integration path. If you do not ship a spec, you are invisible to that loop. The agent has no source of truth to read from; the engineer falls back to guessing path shapes from your docs page; the result is a half-broken client that gets thrown away. Worse: if your competitor’s spec is there, the agent silently picks the path of least resistance, and you lose the integration before the engineer ever read your homepage. This is the shape of distribution in 2026. Specs are the new SDKs. ## What our spec contains `https://.originchain.ai/openapi.json` and the canonical mirror at `https://originchaindb.com/openapi.json` are vanilla OpenAPI 3.1 - no extensions, no vendor-specific quirks, no auth flow that needs explanation. It is the simplest possible shape for a generator to consume. Auth. One scheme: `bearerAuth`, HTTP `Authorization: Bearer `. Every protected path declares it. There is no OAuth flow, no API-key header alternative, no session cookie. One auth path means generators produce one auth method, which means the agent doesn’t have to decide between three. Paths. Six of them: one per query shape, plus schema introspection. ```plaintext POST /v1/tenants/{tenant}/sql Typed SQL with bind params POST /v1/tenants/{tenant}/ask Natural-language question → rows POST /v1/tenants/{tenant}/vector/search HNSW top-k POST /v1/tenants/{tenant}/fts/search BM25 full-text POST /v1/tenants/{tenant}/graph/dijkstra Weighted shortest path GET /v1/tenants/{tenant}/schemas Tenant table introspection ``` Response shapes. Every successful response is a `{ rows: [...], meta: {... } }` envelope. Every error is a `{ error: { code, message, details } }` envelope. The envelope is a generated schema, not a freehand example, so the generator emits a `Response` wrapper type the SDK user can pattern-match on. Examples. Each path has at least one `examples` block on the request body and one on the 200 response. Generators that consume examples (most of the modern ones do) lift them straight into the SDK as fixture tests. ## A 10-line example Feed the spec to a coding agent in a fresh session: ```plaintext Here is an OpenAPI spec: https://originchaindb.com/openapi.json Generate a minimal TypeScript client with one method per path, typed request and response, bearer-token auth, and a single `OriginChain` class. No external deps beyond fetch. ``` What comes back is roughly: ```ts export class OriginChain { constructor(private base: string, private token: string) {} private async post(path: string, body: unknown): Promise { const r = await fetch(`${this.base}${path}`, { method: "POST", headers: { "Authorization": `Bearer ${this.token}`, "Content-Type": "application/json", }, body: JSON.stringify(body), }); if (!r.ok) throw new Error(`${r.status}: ${await r.text}`); return r.json as Promise; } sql(req: { sql: string; params?: unknown[] }) { return this.post<{ rows: T[]; meta: object }>("/v1/tenants/{tenant}/sql", req); } ask(req: { question: string }) { return this.post<{ rows: T[]; meta: object }>("/v1/tenants/{tenant}/ask", req); } vectorSearch(req: { table: string; vector: number[]; k: number; metric?: "cosine" | "dot" | "l2"; mode?: "fast" | "high_recall"; filter?: Record; }) { return this.post<{ rows: T[]; meta: object }>("/v1/tenants/{tenant}/vector/search", req); } } ``` Thirty seconds, no human in the loop, working client. The whole point of shipping the spec is making this trajectory the default. ## What a static spec misses Three things a vanilla OpenAPI spec cannot do, called out so the claim sheet is honest: - Runtime introspection. The spec describes `GET /v1/tenants/{tenant}/schemas` but does not contain your tenant’s actual tables. The agent generating a client knows about the shape of the introspection call but not about your `products` table specifically. For that the agent calls `/v1/tenants/{tenant}/schemas` at runtime - which is the right answer (data shouldn’t live in the spec), but worth saying out loud. - Examples that match real tenant data. The `examples` blocks in the spec are generic - `"sql": "SELECT * FROM products LIMIT 10"`. They are not your products. A generator that wants to produce fixture-realistic tests has to pull a sample from the tenant after auth, and most generators don’t go that far. - Type precision around the JSON-Plan trees. OriginChainDB returns the executed plan tree in the `meta.plan` field for SQL responses, and the tree is recursive: `Scan` and `Filter` and `Project` and `HashJoin` are all variants of a `PlanNode` discriminated union. OpenAPI 3.1 supports `oneOf` discriminators, which we use, but most generators flatten the union to `unknown` rather than emit a tagged sum type. The user’s SDK ends up correct on the wire but loose at the type level. We accept this - the alternative is shipping a custom code generator, which is the kind of vendor lock-in we work hard to avoid. None of these are reasons not to ship the spec. They are the reasons to ship the spec and the language SDKs we hand-write for the precision cases (TypeScript, Python, Go) - the spec is the floor, the SDK is the ceiling, and the agent loop fills in between them. ## What this is NOT The spec at `/openapi.json` is the public surface only. It does not document: - The internal management API for tenant provisioning (which is operator-only). - The replication protocol between primary and follower (which is private wire format). - The `/admin/*` endpoints (which are gated by a separate bearer scope). We deliberately keep the spec narrow. Every endpoint in the spec is one we commit to maintaining the shape of across a major version. If it is in `/openapi.json`, it is part of the contract; if it is not, it is implementation detail and may move. ## Try it The spec lives at [/openapi.json](https://originchaindb.com/openapi.json). The hand-written SDKs and full reference docs live at [/docs/api](https://originchaindb.com/docs/api). Both are stable surfaces; both are versioned in lockstep with the engine. --- # Per-key TTL: agent memory that forgets - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/per-key-ttl-agent-memory Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal ttlagent-memoryephemeraldesigntutorial # Per-key TTL: agent memory that forgets Per-key TTL lets each record expire on its own: a ttl in seconds on any write, invisible to reads at expiry, compacted in the background. Maximum 365 days. OriginChainDB Team May 5, 2026 About 5 min read Agent systems / Retention ## Give short-lived memory an expiry. 1. Write Set the record lifetime 2. Read Use it while it is eligible 3. Expire Reach the configured deadline 4. Reclaim Follow storage cleanup semantics Expiry and physical storage cleanup are distinct events. Choose retention from the purpose of the memory and follow the endpoint's documented TTL semantics. TL;DR - Every OriginChainDB write can carry a `ttl` in seconds. The key stops appearing in reads the moment it expires, and the substrate compacts it away on its own schedule - so tool-call traces, session caches and idempotency keys clean themselves up with no sweeper job and no cron. TTL is per-key, not per-shape, and the maximum is 365 days. ## The problem with permanent agent memory If you store every tool call your agent makes - and you should, for retrieval and reproducibility - you accumulate millions of records per active user. Most of them are useless after 24 hours. The reasoning trace from a failed attempt three weeks ago has zero retrieval value today. The naive approach is to write everything and run a daily DELETE job. This works at small scale and breaks at large scale: the DELETE has to scan, lock, and free space, and on a heavily-loaded store this becomes operationally annoying. You end up running it during off-peak hours, watching for it to take 3x longer than expected, and resigning yourself to “the cleaner is broken again.” The right answer is per-record expiry that the substrate enforces transparently. ## How TTL works in OriginChainDB Every write can carry a `ttl` (time-to-live) in seconds. The key becomes invisible to reads after the TTL expires; the substrate compacts it away on its own schedule. ```typescript // Tool-call trace, expires in 1 hour await oc.put({ shape: "tool-call", id: callId, set: { tool: "search", args, result, latency_ms }, ttl: 3600, }); // User session cache, expires in 30 min await oc.put({ shape: "session-cache", id: sessionId, set: { state }, ttl: 1800, }); // Embedding refresh marker, expires in 24h await oc.put({ shape: "stale-marker", id: entityId, set: { triggered_at: now }, ttl: 86400, }); ``` After the TTL, reads return “not found.” No application code needs to check expiry; no cron job runs. The compaction is the substrate’s responsibility. ## What this lets you build Tool-call memory with automatic forgetting. Store every tool call your agent makes with a 24-hour TTL. The retrieval pipeline searches over the recent window only, automatically. No “forgetting” logic in the application. Session-scoped state. Push intermediate reasoning state with a session-length TTL. When the session ends, it cleans itself up. No “did the user explicitly logout” detection needed. Replay protection. Idempotency keys for incoming requests get a 5-minute TTL. The same request retried within that window deduplicates; outside it, the substrate doesn’t keep state forever. Cache layers. OriginChainDB isn’t a cache, but TTL’d shapes work like one for AI-feature output. Store an LLM-generated response with a 1-hour TTL keyed on the input hash; the next identical input gets the cached answer for free. Embedding staleness markers. When an entity changes, write a `stale-marker` shape with a long TTL. A background job sweeps stale markers and re-embeds. If the entity hasn’t changed during the TTL window, the marker auto-expires and you don’t re-embed needlessly. ## TTL semantics, precisely - TTL is per-key, not per-shape. Two records of the same shape can have different TTLs. - The expiry timestamp is set at write time. Updating a record extends the lifetime to “now + new TTL,” not “original-write-time + new TTL.” - TTL of 0 or unset means “no expiry.” This is the default for shapes that didn’t declare TTL behavior. - Expired keys don’t appear in reads or ANN search results. The substrate’s read path filters them, even before compaction has run. - Compaction is asynchronous. Disk space isn’t freed instantly. For workloads where space matters, expect compaction to run within minutes, not seconds. - Indexes are updated atomically. When a TTL’d record expires, its derived indexes go invisible at the same logical time. There’s no “primary expired but the index still points at it” window. ## What TTL doesn’t do It’s not a scheduling primitive. Don’t use TTL to “trigger” a background job after some interval - there’s no callback. If you need that, write a marker and run a sweeper. It’s not a billing-period primitive. TTLs are seconds-from-write, not “expires at end of month.” If you need calendar-aligned expiry, write a marker shape with a precise `expires_at` timestamp and check it in application code. It’s not a delete-on-event primitive. TTL is purely time-based. To delete on a non-time event (user closes account, etc.), call DELETE explicitly. ## Performance characteristics The TTL machinery adds zero overhead to writes - it’s a few extra bytes per record. Reads check the expiry timestamp inline; this costs about as much as a single attribute access (negligible). The cost is on the compaction path: the substrate has to scan and reclaim expired keys eventually. For workloads that write millions of TTL’d records per day, this is real work - but it runs in the background, throttled to avoid impacting foreground throughput. We tune it conservatively by default. ## FAQ ### What is per-key TTL? Per-key TTL is a database feature where each individual record can carry its own expiry timestamp. After the TTL, the record becomes invisible to reads, and the substrate compacts it away in the background. It’s the standard pattern for ephemeral data - caches, sessions, idempotency keys. ### How does this differ from Redis TTL? Redis is purely in-memory; OriginChainDB TTL is on durable storage with the same durability semantics as any other write. You get TTL without giving up persistence guarantees. ### Can I extend a TTL? Yes - the next write to the key resets the expiry to “write-time + TTL.” A common pattern for session refresh. ### What’s the maximum TTL? Bounded by the substrate at 365 days. If you want longer expiry than that, store an explicit timestamp and DELETE on a sweeper. ### Does TTL apply to vector shapes? Yes. A vector with TTL becomes invisible to ANN search after expiry. Useful for ephemeral embeddings of session state. ## What to read next - [OriginChainDB quickstart](https://originchaindb.com/blogs/quickstart) - how to write your first TTL’d record. - [Idempotent tool calls](https://originchaindb.com/blogs/idempotent-tool-calls) - the most common production use of TTL. - [Schemas](https://originchaindb.com/docs/schemas) - the underlying data model. --- # OriginChainDB quickstart in five minutes - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/quickstart Published: 2026-05-04T20:59:15.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T20:18:50.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal tutorialquickstartsdkgetting-startedvector-search # OriginChainDB quickstart in five minutes Provision a managed instance, write a JSON record and its vector embedding in one atomic transaction, run a similarity search, hydrate the match. OriginChainDB Team May 4, 2026 About 5 min read Getting started / First workflow ## From a schema to a useful query. 1. Connect Configure endpoint and credentials 2. Model Register fields and a primary key 3. Write Submit the required representations 4. Query Read results and their source IDs Start with a small record and a query you can inspect. Follow current quickstarts for endpoint details and separate vector or full-text ingestion steps. TL;DR - Provision an OriginChainDB instance from the dashboard, install the SDK, then write a JSON payload and its vector embedding in one atomic transaction and read the record back by similarity search. No infrastructure, no schema migration, no glue code - the shape declaration is the schema. ## What you’ll need - An [originchaindb.com](https://originchaindb.com/) account (free database; no credit card required). - Node.js 18+ or Python 3.10+ on your local machine. - Five minutes. That’s it. No Docker, no provisioning scripts, no CLI auth dance. ## Step 1 - Provision an instance Sign in at [originchaindb.com/login](https://originchaindb.com/login), click + New instance on the dashboard, name it (e.g. `dev-1`), and pick a region. The smallest configuration is what the free database gives you - adequate for everything in this guide. Provisioning takes under two minutes. When the status flips to running, you’ll see the instance card show an HTTPS endpoint that looks like: ```plaintext https://abc123def.originchain.ai ``` Copy that. It’s your tenant endpoint. Click the instance to drill in, then API keys → + New key. Copy the key (shown once). This is your bearer token for everything below. ## Step 2 - Install the SDK ```bash # Node / TypeScript npm install @originchain/sdk # Python pip install originchain ``` Both clients are thin wrappers over the HTTP API - if your stack isn’t covered, the [OpenAPI spec](https://api.originchain.ai/openapi.json) is auto-generated and works with any client generator. ## Step 3 - Connect ```typescript import { OriginChainClient } from "@originchain/sdk"; const oc = new OriginChainClient({ baseUrl: process.env.OC_ENDPOINT!, // from step 1 bearer: process.env.OC_API_KEY!, // from step 1 }); await oc.health; // → { ok: true } ``` ```python from originchain import OriginChain oc = OriginChain( endpoint=os.environ["OC_ENDPOINT"], api_key=os.environ["OC_API_KEY"], ) oc.health # → {"ok": True} ``` ## Step 4 - Declare a shape A key shape describes how a logical entity maps onto the substrate. We’re modelling articles with text + a vector embedding: ```typescript await oc.shapes.create({ shape: "article", key: "article/{id}", value: { title: { type: "string", indexed: true }, body: { type: "string" }, author: { type: "string" }, }, }); await oc.shapes.create({ shape: "article-embedding", key: "vec/article/{id}/body", value: { type: "f32", dim: 768 }, }); ``` That’s it. No DDL, no migration. The shape is the schema. (For more on shapes, see [Schemas](https://originchaindb.com/docs/schemas).) ## Step 5 - Write a record ```typescript const article = { id: "intro-to-rag", title: "Introduction to RAG", body: "Retrieval-augmented generation is a technique for...", author: "Lalith", }; const embedding = await yourEmbeddingModel(article.body); // 768-dim f32 await oc.transaction([ { shape: "article", id: article.id, set: article }, { shape: "article-embedding", id: article.id, set: embedding }, ]); ``` Both writes - the JSON record and the vector - commit in the same atomic operation. Either both make it durable or neither does. That’s the point of one substrate: every shape sees the row, or none do. (For more on the atomicity guarantee, see [Transactions](https://originchaindb.com/docs/transactions).) ## Step 6 - Read it back Direct read by id: ```typescript const a = await oc.get("article", "intro-to-rag"); // → { id: "intro-to-rag", title: "...", body: "...", author: "Lalith" } ``` Lookup by indexed attribute: ```typescript const found = await oc.findOne("article", { title: "Introduction to RAG" }); // → same article ``` Vector similarity search: ```typescript const queryVec = await yourEmbeddingModel("how does retrieval-augmented generation work"); const matches = await oc.vectorSearch("article-embedding", { query: queryVec, top_k: 5, }); // → [{ id: "intro-to-rag", distance: 0.18 },...] // Hydrate matched articles const articles = await Promise.all( matches.map((m) => oc.get("article", m.id)) ); ``` That’s the loop. JSON payload + vector + similarity search, all atomic, all from one client. ## Step 7 - Wire it into your app Most AI features fit one of three patterns at this point: Personalized retrieval - embed user queries, ANN-search a content corpus, hydrate the top hits. Steps 5-6 are the loop; the rest is your prompt-engineering taste. Tool-call memory - every tool call your agent makes writes a record + a vector. Future calls can search the past for “have we tried this before?”. Live feature stores - every user event writes a record; your model reads the most recent N at inference time. Last-writer-wins handles concurrent updates correctly without retry loops. ## Going to production When you flip from the free database to a paid plan, three things change: - Compute plan scales the writer node’s CPU + memory. - Storage plan sets the durable footprint cap. - HA replicas become available - the snapshot-bootstrap design means a follower can take over on writer failure without the seeding gap a naive replica copy leaves behind - though replication is asynchronous, so an abrupt primary loss can still cost the most recent acknowledged writes. (See [HA snapshot bootstrap](https://originchaindb.com/blogs/ha-snapshot-bootstrap).) There’s no migration step. The instance you started on the free database is the instance you keep. ## FAQ ### How fast is this in practice? Single-record reads are sub-millisecond from a co-located client. Writes are ~280 µs at the 50th percentile under sustained load. Vector search latency depends on corpus size - around 5 ms at 10M vectors with default ANN settings. ### Do I need to manage the database? No. OriginChainDB is managed-cloud only. You don’t see the underlying compute, you don’t run the upgrades, you don’t tune storage internals. You write records and read them back. ### What languages have official SDKs? TypeScript/JavaScript and Python today. Go and Rust are next on the SDK roadmap. The HTTP API is documented as OpenAPI so any code generator works in the meantime. ### Can I run this locally for dev? The substrate is a managed service; there’s no self-hosted binary. For local dev, hit the free database - it costs nothing, and dev tenants are sized for $0.50/day once you outgrow it. ### What if I need to change the shape? Add a field to a value-side schema and old records keep working with the new field defaulting to null. Renames and type changes are version-bumped: declare a new shape version, the substrate runs a dual-read transform during cutover, no downtime. ## What to read next - [Schemas](https://originchaindb.com/docs/schemas) - the data model in detail. - [Transactions](https://originchaindb.com/docs/transactions) - why the multi-shape write in step 5 is atomic. - [Troubleshooting](https://originchaindb.com/docs/troubleshooting) - common problems and fixes when a call doesn’t behave. - [One database, every query shape](https://originchaindb.com/blogs/one-database-every-query-shape) - the bigger architectural picture. --- # RAG at 10M documents: version skew - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/rag-at-10m-documents-the-four-vendor-failure-mode Published: 2026-05-26T08:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal ragvector-searchhybrid-searchengineering # RAG at 10M documents: version skew In a four-database RAG stack the vector, BM25 posting and row text of one chunk can commit at different versions - and the faithfulness check grades the skew. OriginChainDB May 26, 2026 About 7 min read Retrieval / Version alignment ## Keep retrieved evidence tied to its source version. 1. Source text Passage ID + document version 2. Vector entry Embedding for that passage 3. Text index Terms from the same content 4. Answer context Validate IDs before generation A shared passage ID and version make mismatches detectable. The application must coordinate ingestion and decide when updated representations are ready to query. In a RAG stack built on four databases, a single chunk’s embedding, BM25 posting and row text can end up committed at three different versions, because nothing spans the four systems transactionally. The Stage 7 faithfulness check then grades the answer against three versions of the same chunk and returns a number that no longer means what you think it means. On a single-substrate engine the skew is not reachable, because every shape of the chunk commits or none does. A widely-shared Medium piece walks through a production RAG architecture for 10 million documents. The 10 stages - chunk, embed, hybrid retrieve, ANN plus rerank, confidence gate, constrained generate, cite, verify, semantic cache, trace - are the right blueprint. The article gets the retrieval math right, the confidence thresholds right, the faithfulness check right. We have nothing to add to that part. What it doesn’t discuss is the failure mode the stack itself introduces. Almost every published RAG architecture at this scale assumes four separate systems: - A vector database (Pinecone, Weaviate, pgvector) for embeddings - A search engine (Elasticsearch, OpenSearch, Vespa) for BM25 - A relational store (Postgres) for chunk metadata and source provenance - A cache (Redis, sometimes a second small vector index) for semantic deduplication Four databases. Every document write must succeed against all four. Every document update must invalidate all four. And nothing in those four systems is jointly transactional with the others. This is fine when you have 10,000 documents and writes are infrequent. It becomes a class of silent bug at 10 million. ## The bug A user updates a regulatory filing. Your ingest pipeline re-chunks it, re-embeds each chunk, and writes: 1. Pinecone gets the new embedding for chunk `doc-471:12`. 2. Elasticsearch gets the new BM25 posting for the same chunk. 3. Postgres gets the new row text and version timestamp. Three writes, three systems. They do not commit atomically. There is no shared transaction. There is no shared commit boundary. In the normal case nothing breaks - the three writes finish in milliseconds and the user moves on. But the moment one fails, retries, or even just lands out of order under load, you have an internally inconsistent corpus. A query against that corpus can see: - Pinecone returning chunk `doc-471:12` at vector version `v3` - Elasticsearch returning the same chunk at posting version `v2` - Postgres metadata showing the chunk is at row version `v1` Your retriever picks the chunk. Your reranker scores it. Your LLM generates an answer using the v1 text. Your Stage 7 faithfulness checker - the one that verifies every assertion in the answer is grounded in the retrieved chunks - sees the v3 vector tag, the v2 BM25 posting, and the v1 text. The faithfulness math runs against three versions of “the same chunk” and produces a number that no longer means what you think it means. That number then drives a decision: do we show the answer, do we re-run with a stricter prompt, do we escalate to a human? You are making a quality-gate decision based on a corpus that is, briefly, lying to you. In our experience this happens enough that it is the dominant source of unexplained hallucination reports at scale. It is never on the dashboard because no system reports the skew - each individual database is healthy. It is never in the logs because the writes all succeeded. It surfaces as an LLM that “sometimes makes things up”, and the postmortem inevitably ends with someone saying “we tightened the prompt”. ## The fix is not at the application layer The natural reaction is to add cross-system reconciliation: write Postgres first, then Pinecone, then Elasticsearch, all behind a saga, with a worker that detects skew and re-syncs. We have seen teams build elaborate machinery for this. It works. Then it breaks under load. Then the worker itself becomes a thing to monitor. Then somebody starts skipping a step “because it is slow” and three weeks later you have skew again. The fix is not at the application layer because the consistency boundary belongs to the engine, not the app. If your vectors, postings, and rows live in different systems addressed by different bearers and persisted separately, no amount of application code can give you atomic cross-shape writes. You can get eventual consistency, but Stage 7 of the canonical RAG pipeline asks for point-in-time consistency. ## What we built OriginChainDB is a single-substrate database where SQL rows, vector indexes, full-text indexes, and graph edges share one managed k/v store and one commit boundary. The write path for a document chunk is a single atomic operation that touches every shape at once: ```ts await oc.transaction([ { shape: "rows", table, id, set: row_payload }, { shape: "vec", table, id, set: embedding_payload }, { shape: "fts", table, field, id, set: posting_payload }, ]); ``` One atomic write — all shapes or none. After a crash either every shape applies or none do. There is no path through the system where the vector lands and the row doesn’t, or where the BM25 posting commits and the vector doesn’t. The Stage 7 faithfulness check then runs against a corpus that is internally consistent by construction. The version-skew failure mode doesn’t appear in postmortems because it isn’t reachable. ## What the application still owns A single-substrate engine is necessary but not sufficient for great RAG. The article’s other nine stages still apply, and they are still your code: - Chunker - the right window size depends on your domain. We don’t ship one. - Embedder - `text-embedding-3-small`, Cohere `embed-v3`, a managed embedding model, or a local ONNX model. Your choice, your bill. - Reranker - Cohere Rerank, Voyage, or a local cross-encoder. Only worth wiring at >1M chunks. - Constrained-generation prompt - temperature 0.0, citation-mandatory, “say I don’t know” prefix. - Faithfulness verifier - extract assertions with NER + regex, ground each in retrieved chunks, fall back at <0.8. What you stop owning is the coordination problem. The atomic-write guarantee lets the verifier compute a number that actually means something. ## What the code looks like Ingest, in TypeScript with the SDK: ```ts for (let i = 0; i < chunks.length; i++) { const id = `${docId}:${i}`; await oc.sql( `INSERT INTO chunks (id, source, page, text, created_at) VALUES (?, ?, ?, ?, ?)`, [id, source, chunks[i].page, chunks[i].text, Date.now], ); await oc.vectorPut("chunks", { id, embedding: emb.data[i].embedding, dim: 1536, metric: "cosine", metadata: { source, page: chunks[i].page }, }); await oc.ftsIndex("chunks", "text", { doc_id: id, text: chunks[i].text }); } ``` Read, in parallel: ```ts const [vecHits, ftsHits] = await Promise.all([ oc.vectorTopk("chunks", { query: qVec, k: 20, dim: 1536, metric: "cosine" }), oc.ftsSearch("chunks", "text", { q: question, mode: "bm25", k: 20 }), ]); const top = reciprocalRankFusion(vecHits, ftsHits).slice(0, 5); const rows = await oc.sql( `SELECT id, source, page, text FROM chunks WHERE id IN (?, ?, ?, ?, ?)`, top.map(h => h.id), ); ``` The fuller walkthrough - including the semantic-cache pattern, configuration guidance for corpus sizes from 100K to 10M chunks, and the boundary diagram - lives in [the RAG docs page](https://originchaindb.com/docs/rag). ## On scale, honestly The Medium article uses 10 million documents as its target. On OriginChainDB, 10 million 1536-dimensional chunks is roughly: - 60 GB of embeddings - 30 GB of BM25 postings - ~10 GB of row text and metadata That is a mid-size configuration - a single dedicated instance with enough RAM to keep the HNSW graph hot. HNSW recall@10 above 0.95 holds at that scale; query latency stays in the tens of milliseconds before the reranker. Beyond 50 million chunks you are into multi-writer sharding territory, which is on our depth-first roadmap to 1.0 but not GA yet. Talk to us if you are already there. ## The blueprint, with three fewer databases The Medium article’s blueprint works. Removing three of its vendors (Pinecone, Elasticsearch, Redis-vector) removes the cross-system version skew with them; the 10-stage architecture and the retrieval math are unchanged. If you are building RAG at any scale where “we sometimes hallucinate” is a thing you are tracking, the question worth asking is not “what is my reranker doing” - it is “what is my corpus saying when I read it three different ways at the same instant”. --- # RAG latency budget: where the time goes - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/rag-latency-budget Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal ragperformancelatencyanntutorial # RAG latency budget: where the time goes A 2-second conversational RAG loop spends 500-2000ms in the LLM call; embed, ANN, hydrate and prompt should fit under 250ms combined. OriginChainDB Team May 5, 2026 About 6 min read Performance / Retrieval pipeline ## Measure each stage of the answer. 1. Embed Prepare the query vector 2. Retrieve Search the relevant indexes 3. Hydrate Load source records and context 4. Generate Call the model and stream output End-to-end latency includes the database and the surrounding application. This is a sequence diagram, not a proportional timing chart. TL;DR - A conversational RAG loop has a 1-3 second user-perceived budget, and 500-2000ms of it is the LLM call. Embedding, ANN search, hydration and prompt construction should sit under 250ms combined. When they don’t, hydration is usually the reason: serial GETs instead of a parallel batch or an `mget`. ## The budget Pick your latency target first. Common targets: - Conversational AI: <2s end-to-end. Anything longer feels broken. - Search-with-AI-summary: <3s. People expect search to be slow-ish. - Background agent: <30s. Latency doesn’t matter; reliability does. - Streaming output: <800ms to first token. The rest streams. This post focuses on the conversational-AI target: 2 seconds end-to-end. ## Where the time goes A typical RAG request has 5 stages: 1. Embed the query → ~50-200ms (depends on the model + provider) 2. ANN search → 1-50ms (depends on corpus size + index quality) 3. Hydrate matched records → 1-100ms (depends on storage + parallelism) 4. Construct the prompt → ~5-20ms (string ops + token counting) 5. LLM generation → 500-2000ms (this is most of your budget) Add latency variance from network round-trips, cold caches, and the occasional 99th-percentile outlier. The LLM is your bottleneck. Stage 5 is the hard floor. Everything else combined should sit under 250ms or you’re losing real time. ## Realistic ANN search numbers | Corpus size | OriginChainDB p50 | OriginChainDB p99 | | --- | --- | --- | | 10k vectors | <1ms | ~3ms | | 100k vectors | ~1ms | ~5ms | | 1M vectors | ~2ms | ~8ms | | 10M vectors | ~5ms | ~15ms | | 100M vectors | (post-optimiser) | (post-optimiser) | These are end-to-end from the SDK call to the result, including the HTTP round-trip from a co-located client. Cross-region adds whatever your network latency is. What this tells you: ANN itself is rarely the bottleneck below ~10M vectors. If your search is taking 100ms, the time is going somewhere else (network, hydration, embedding regeneration). ## The hydration step Most RAG implementations get back a list of ids from ANN search and then need to fetch the actual content. This is the step that’s surprisingly easy to screw up. Wrong: serial GETs. ```typescript const ids = await oc.vectorSearch("article-embedding", { query: q, top_k: 10 }); const articles = []; for (const m of ids) { articles.push(await oc.get("article", m.id)); // 10 round-trips! } // → 50-100ms wasted on round-trip overhead ``` Right: parallel batch. ```typescript const ids = await oc.vectorSearch("article-embedding", { query: q, top_k: 10 }); const articles = await Promise.all( ids.map((m) => oc.get("article", m.id)) ); // → 1 round-trip's worth of latency, ~3-5ms ``` Even better when supported: `mget`. ```typescript const ids = await oc.vectorSearch("article-embedding", { query: q, top_k: 10 }); const articles = await oc.mget("article", ids.map((m) => m.id)); // → 1 actual round-trip + server-side parallelism, ~2-3ms ``` That’s a 30x improvement on the hydration step from one syntax change. We’ve seen real applications discover this and find half their latency budget by fixing the loop. ## The embedding step Embedding the query is often outside your database - you’re calling OpenAI or an internal model. This is the easiest place to optimize: - Cache aggressively. If the same query was embedded recently, reuse the embedding. A 5-minute cache on common queries is free latency. - Pre-compute embeddings for known structured queries (FAQ entries, common templates). - If your model supports it, use a smaller embedding model for the query and a larger one for the corpus. Asymmetric models exist. A 100ms embedding call you can avoid is worth 100ms of any other optimization. ## The prompt construction step Looks innocent, often isn’t. Watch for: - Token counting on long contexts. If you’re trimming retrieved content to fit a model’s context window, the trimming logic should be O(n) on the trimmed length, not the original. - String concatenation on long context. Use array `join` instead of `+=` accumulation. - Reserialization. If the retrieved record is already a JSON string, don’t re-`JSON.stringify` it. These look like 5ms problems each but they compound when you’re concatenating dozens of retrieved records. ## The LLM step Mostly out of your control, but a few levers: - Stream the response. Time-to-first-token is what users perceive; total time is less important. - Pick the smallest sufficient model. - Use prompt caching if your provider supports it. The system prompt + retrieved-context structure is often cacheable across requests. - Parallelize tool calls if your loop can. Independent tool invocations are a tree, not a list. ## Measuring your own loop The single most useful thing you can do: instrument every stage with `Date.now` (or your tracer of choice) and log the per-stage durations on every request. After 100 requests, your bottleneck is obvious. ```typescript const t0 = Date.now; const queryEmbedding = await embed(query); const t1 = Date.now; const matches = await oc.vectorSearch("article-embedding", { query: queryEmbedding, top_k: 10 }); const t2 = Date.now; const articles = await oc.mget("article", matches.map((m) => m.id)); const t3 = Date.now; const prompt = buildPrompt(query, articles); const t4 = Date.now; const response = await llm.complete(prompt); const t5 = Date.now; logStageTiming({ embed_ms: t1 - t0, ann_ms: t2 - t1, hydrate_ms: t3 - t2, prompt_ms: t4 - t3, llm_ms: t5 - t4, }); ``` Plot the percentiles. Find the stage with the highest p99. Fix that one. Repeat. ## A worked budget For a 2-second conversational target, our default budget allocation: | Stage | Target p99 | What if it’s higher | | --- | --- | --- | | Embed | 200ms | Cache common queries; smaller model for queries | | ANN | 20ms | Corpus is too large; consider filter-then-search | | Hydrate | 10ms | You’re not parallelizing or batching | | Prompt | 20ms | Token counting / serialization issue | | LLM | 1500ms | Smaller model; streaming; prompt caching | | Network/overhead | 250ms | Move client closer to substrate | If your numbers look very different from this, it tells you where to look first. ## FAQ ### How fast is OriginChainDB ANN at production scale? About 1ms p50 at 100k vectors, ~2ms at 1M and ~5ms at 10M, measured end-to-end from a co-located client. 100M is post-optimiser work. Full table above. ### Should I co-locate my application with OriginChainDB? Yes. A 50ms cross-region RTT on every database call eats your budget fast. Run your application in the same region as your tenant. ### Is hydration really 30x faster as parallel batch? Yes, in real loops we’ve seen. The fix is ~3 lines of code. ### Does caching the embedding break personalization? If the cache key is the raw query string, no - same query string produces the same embedding regardless of user. If you want per-user embeddings, key the cache by `(user, query)`. ### What about latency variance? The p99 is what your users complain about, not the p50. Always measure both. ## What to read next - [OriginChainDB quickstart](https://originchaindb.com/blogs/quickstart) - the basic loop you’d be measuring. - [Benchmarks](https://originchaindb.com/benchmarks) - the throughput and latency runs we have published. - [Why we don’t need a separate vector database](https://originchaindb.com/blogs/no-separate-vector-database) - eliminating the dual-write hop saves a round-trip. --- # The token economy and your AI bill - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/token-economy-occi Published: 2026-07-08T09:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T20:18:50.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal ai-coststoken-economyoccicontextual-intelligenceenterprise # The token economy and your AI bill Token spend is now a governed line item, and its biggest driver is missing context: an under-informed model retries, over-prompts and calls other models. Zaheer Jul 8, 2026 About 5 min read AI applications / Context ## Measure the context that reaches the model. 1. Retrieve Find relevant source material 2. Select Keep evidence within the budget 3. Generate Build and run the prompt 4. Evaluate Review quality and token usage Context selection, answer quality and token consumption should be evaluated together. The diagram illustrates an application workflow, not automatic cost savings. Token spend has become a governed line item rather than a technical detail, and the largest driver of it is not model size — it is missing context. An under-informed model compensates by generating longer prompts, retrying, looping, and calling other models to fill the gap, and every one of those behaviours is billable. For two years, the AI industry celebrated consumption. More tokens, more agents, more autonomous loops — every metric pointed up and to the right, and nobody was counting the cost. That era is ending. We are entering the token economy, where the price of every inference is a line item, and where the organizations that win will be the ones that extract the most intelligence per dollar. ## Token consumption has quietly become one of your largest hidden costs AI token usage has crossed the threshold from a technical detail to a strategic cost center. “Token maxing” — inefficient or runaway token consumption — can generate enormous operating expense at scale, and the numbers are no longer hypothetical. Consider the scale being reported across the industry: - Meta is reported to consume on the order of 73 trillion tokens per month across its workforce — which, at prevailing enterprise pricing, translates to well over $220 million per month. - Uber reportedly burned through its entire annual AI budget in just four months. - Heavy individual users are spending close to $35,000 per seat, launching workflows, running coding agents, executing autonomous loops, and delegating tasks to models that call other models. At that level, consumption is essentially machine-generated — not human-paced. The distribution is even more revealing. The median employee spends around $12 per month. The top 10% spend in the hundreds. The top 1% spend in the thousands. A tiny minority spends in the tens of thousands. In other words: a handful of automated, high-intensity workloads can quietly dominate your entire AI bill. ## ”Token usage is not productivity” That line, increasingly heard from AI leaders, captures the turning point. For a while, high consumption was treated as a proxy for value — a signal that teams were “adopting AI.” Leaders now recognize that burning tokens is not the same as creating outcomes. The parallel is exact. Cloud spend and SaaS spend both went through this maturation: an exuberant experimentation phase, followed by the arrival of governance, budgets, monitoring, and finance-team scrutiny. AI spend is next. Finance teams are already introducing controls, budgets, and dashboards to track token usage. And going forward, the cost curve itself will determine the pace of adoption. If AI unit economics don’t work, the transformation stalls — regardless of how good the models are. The strategic implication for every enterprise and startup deploying AI: model quality is table stakes; cost-aware architecture is the competitive advantage. Efficiency is no longer an engineering nicety. It is a boardroom concern. ## The problem isn’t the model. It’s the context. Here’s what most organizations miss. The vast majority of wasted tokens don’t come from the model being too large — they come from the model being under-informed. When an LLM lacks the right context, it compensates: it generates longer prompts, retries, hallucinates, gets corrected, loops, and calls other models to fill the gap. Every one of those compensating behaviors is billable. Poor context is expensive twice over. You pay for the wasted tokens, and you pay for the bad outcomes — the hallucinations, the low-confidence decisions, the manual rework, the eroded trust. The answer is not simply “use a cheaper model” or “cap everyone’s budget.” The answer is to make every token count by feeding the model richer, more deterministic context — so it gets to the right answer faster, with fewer attempts, less sampling, and less hallucination. That is precisely what OriginChainDB built. ## Introducing OCCI: OriginChainDB Contextual Intelligence OCCI (OriginChainDB Contextual Intelligence) is a context intelligence layer that sits between your data and your models. Rather than treating tokens as an unavoidable tax, OCCI treats context as the lever. Organizations deploying OCCI reduce AI token consumption by 30–50% while improving business outcomes. With OCCI, you can: - Build a richer contextual intelligence layer for more deterministic, repeatable AI outputs — fewer retries, fewer wasted tokens. - Reduce hallucinations and improve response accuracy, so your teams trust the output and stop paying for rework. - Enhance risk intelligence and decisioning, turning raw model output into governed, defensible business decisions. - Generate better Next Best Actions (NBA) that move revenue and retention, not just generate text. - Govern LLM outputs through optimized sampling parameters, giving you control over the cost/quality tradeoff instead of leaving it to chance. - Reduce latency and improve application performance, so faster responses cost you less, not more. - Lower AI operating costs significantly — turning your token line item from an open-ended liability into a managed, predictable investment. ## What this means for the people who own the number For the CIO: OCCI brings governance, determinism, and performance to an AI stack that has been running without guardrails. You get architecture that scales without your cost curve scaling with it — and outputs your organization can actually rely on. For the Head of AI: Better accuracy, fewer hallucinations, and stronger unit economics on every application you ship. OCCI lets your engineers and product teams build ambitious AI without every experiment quietly detonating the budget. For the CFO: AI spend is about to get the same scrutiny as cloud and SaaS. OCCI gives you a lever to cut 30–50% of token cost while improving the quality of what you’re paying for — the rare optimization that doesn’t force a tradeoff between cost and value. For the CEO: In the token economy, efficiency is a moat. The companies that master intelligence-per-dollar will out-ship and out-scale competitors who are still celebrating consumption. OCCI turns cost discipline into a durable competitive advantage. ## The experimentation phase is over. The optimization phase has begun. Whether you’re an enterprise scaling GenAI across the organization or a startup obsessing over AI unit economics, the mandate is the same: build and use AI economically. --- ### Reduce AI costs. Improve AI outcomes. OriginChainDB Contextual Intelligence (OCCI) — more intelligence per dollar. Visit [originchaindb.com](https://originchaindb.com/) or reach us at [zaheer@originchain.ai](mailto:zaheer@originchain.ai). Is token maxing driving up your AI costs? Let’s talk about what a smarter context layer could return to your bottom line. --- # OriginChainDB vs DynamoDB: when each fits - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/vs-dynamodb Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal dynamodbcomparisonkey-valuearchitecturevector-database # OriginChainDB vs DynamoDB: when each fits DynamoDB stores vectors as opaque attributes and needs a sidecar ANN service; OriginChainDB commits the entity and its vector in one transaction. OriginChainDB Team May 5, 2026 About 5 min read Architecture choices / Access patterns ## Start with the queries the application needs. 1. Key access Identify point-read patterns 2. Capacity Model traffic and operations 3. Other models List search and relationship needs 4. Integration Plan ingestion and result assembly Compare access patterns, operating model and integration work. Validate current features and pricing when making a product decision. TL;DR - DynamoDB is a brilliant operational substrate for KV workloads - flat fees, infinite scale, no servers. OriginChainDB is the right pick when your workload is AI-shaped: vectors as first-class shapes, atomic multi-shape writes, sub-millisecond reads with declarative indexes. DynamoDB is the more operationally mature of the two; OriginChainDB removes the second service. ## What each is for DynamoDB is a managed key-value/document store with single-digit-millisecond p99 reads, flat-fee scaling, and zero operational surface. Its data model is partition key + sort key + attributes; you query by exact key match, by sort-key range, or by Global Secondary Index (GSI). Vectors are not first-class - you’d store them as binary blobs and use a sidecar service for ANN. OriginChainDB is a managed AI-native database with managed engine, declarative key shapes, atomic multi-shape transactions, and built-in vector search. Reads are sub-millisecond, writes are batched for high throughput, and the schema is evolved by adding shapes rather than running migrations. They overlap on: KV access patterns, managed-only operation, managed engine, sub-ms reads. They diverge on: vector support, multi-shape transactions, schema evolution, query model. ## When to pick DynamoDB - Your workload is keyed lookups + range scans on a sort key, with a stable shape. - You’re already on AWS and the IAM/CloudWatch integration is non-negotiable. - Vectors aren’t part of the picture, or are handled by a separate service (OpenSearch, Pinecone, etc.). - You’re comfortable with eventual consistency on GSIs (default) or willing to pay for strongly-consistent reads. - Your access patterns are baked - DynamoDB punishes pattern changes after launch. ## When to pick OriginChainDB - Your application stores vectors next to the entity they describe and you want them to be atomic with the entity write. - Your writes are agent-paced (hundreds per session, thousands of sessions). DynamoDB at 1,000 WCU/sec costs real money; OriginChainDB batches many concurrent writers into each durable commit. - Your schema is moving. DynamoDB’s “no schema” is actually “every reader handles every shape variant” - that pain shows up later. OriginChainDB enforces shape contracts at write time so the chaos doesn’t accumulate. - You want a typed query API that knows about your shapes, not raw KV operations. ## The vector story This is where the gap is widest. In DynamoDB, you’d: 1. Store the vector as a binary attribute on the item. 2. Run a separate ANN service (OpenSearch with `knn`, Pinecone, Weaviate, etc.). 3. Stream item changes from DynamoDB Streams to keep the ANN index in sync. 4. Live with the eventual-consistency window: there’s a stretch where the item exists but the vector index doesn’t see it yet. In OriginChainDB: 1. Declare a vector shape next to the entity shape. 2. Write the entity + vector in one transaction. They commit in the same atomic operation. 3. ANN search is just another query against the vector shape. 4. No second service. No stream. No replication lag. For workloads where “the entity and its vector must be visible together” matters (recommendation, personalization, fresh content surfacing), the OriginChainDB model removes a class of bugs. ## The cost comparison Honest numbers, since this is what most people actually care about. | Scenario | DynamoDB on-demand | OriginChainDB entry configuration | | --- | --- | --- | | 100K reads + 100K writes, 1KB items, per month | ~$15 | <$20 (within free allowance) | | 10M reads + 1M writes per month | ~$150 (reads dominate) | $50-150 (depends on storage) | | 100K writes/day with vectors + ANN search | DynamoDB + Pinecone Pro: ~$700+ | $50-200 | | 1B reads/month, 1KB items | ~$1,200 (consider provisioned) | Talk to us - pricing tiers up | Both are cheap at small scale. The vector + ANN combination is where the OriginChainDB economics get interesting - you save on the second service. ## Operational delta DynamoDB: - Capacity modes: on-demand or provisioned (with auto-scaling). - Point-in-time recovery: 35 days. - Cross-region replication: Global Tables. - Backup: continuous + on-demand. - Schema evolution: just write the new attributes. OriginChainDB: - No capacity mode - the substrate handles its own backpressure. - Point-in-time recovery: opt-in, restores land on a 5 s roll-up boundary, retention configurable up to 35 days. - Cross-region replication: not yet (post-1.0). - Backup: continuous change-log archive. - Schema evolution: declare a new shape version, the substrate dual-reads during cutover. DynamoDB is operationally more mature - it’s been in production for 12+ years. OriginChainDB is at the “depth-first to 1.0” stage. Be honest about that when picking. ## FAQ ### Is OriginChainDB a DynamoDB replacement? For AI-native workloads, often yes. For broad enterprise KV use cases, not yet - DynamoDB’s operational maturity (Global Tables, full IAM integration, decade of customer-driven features) is hard to match. We’re depth-first to 1.0 - see [the roadmap](https://originchaindb.com/blogs/depth-first-roadmap-to-1-0). ### Does OriginChainDB have a sort key? Not as a first-class concept. The managed engine gives you point lookups in O(1) and prefix scans on encoded keys. If your DynamoDB usage was sort-key-heavy (range scans by timestamp, etc.), the migration takes thought. ### Can I migrate from DynamoDB? Yes. Each table becomes a shape; each GSI becomes a derived shape. Vectors that were sidecar’d to OpenSearch/Pinecone become co-located shapes. We’ve seen 1-2 weeks of engineering time for medium-complexity migrations. ### What about Global Tables / multi-region? Not on the OriginChainDB roadmap before 1.0. If multi-region is a hard requirement today, stay on DynamoDB. ### How does pricing scale at large volumes? DynamoDB on-demand gets expensive past ~1B requests/month - you’d switch to provisioned + auto-scaling. OriginChainDB pricing currently tiers by compute + storage plan; talk to us for enterprise volumes. ## What to read next - [Schemas](https://originchaindb.com/docs/schemas) - the data model that replaces DynamoDB’s flat document model. - [HA snapshot bootstrap](https://originchaindb.com/blogs/ha-snapshot-bootstrap) - how OriginChainDB handles failover compared to DynamoDB Global Tables. - [Benchmarks](https://originchaindb.com/benchmarks) - the throughput and latency runs we have published. --- # OriginChainDB vs Pinecone: when each fits - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/vs-pinecone Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal pineconevector-databasecomparisonarchitectureann # OriginChainDB vs Pinecone: when each fits Pinecone holds the vector and assumes the entity lives elsewhere, so you dual-write. OriginChainDB writes entity and vector in one transaction. Where each fits. OriginChainDB Team May 5, 2026 About 5 min read Architecture choices / Vector search ## A vector query is part of a larger workflow. 1. Vector workload Measure index and query needs 2. Source records Locate the underlying entities 3. Other queries Account for SQL, text and graph 4. Operations Compare ownership and updates Evaluate retrieval quality alongside record storage and operational requirements. The illustration does not rank vendors or assert benchmark superiority. TL;DR - Pinecone is a dedicated vector database - it stores embeddings, runs ANN search, and assumes your entity data lives somewhere else. OriginChainDB is a unified substrate where vectors are a shape alongside JSON entities, atomic with their parent records. If you find yourself dual-writing “the entity” and “the vector” today, OriginChainDB removes that whole layer. ## What each is for Pinecone is a managed vector database. You give it embeddings + a small JSON metadata blob; it gives you back nearest-neighbor search with filters. It’s purpose-built, well-tuned, and the operational story is clean. The constraint is that everything else (the document the vector represents, the user it belongs to, the timestamp it was generated at) lives in some other database that you have to keep in sync. OriginChainDB stores vectors as a key shape co-located with their parent entity. A `user` shape and a `vec/user/{id}/profile` shape live in the same substrate, are written in the same atomic operation, and queried through the same client. The ANN engine reads from the substrate directly - no replication, no sidecar. ## The architecture you don’t have to build Most teams running Pinecone in production have something like this: ```plaintext [Postgres/Mongo/DynamoDB] <-- source of truth for entities | v [Stream / Worker] <-- watches inserts, generates embeddings | v [Pinecone] <-- ANN-searchable vector index | v [API endpoint] <-- Pinecone returns ids + scores | v [Postgres/Mongo/DynamoDB] <-- hydrate entity by id ``` That’s three round-trips per query, two systems to keep in sync, and a class of bugs around the lag between “entity inserted” and “vector visible to ANN search.” OriginChainDB collapses it: ```plaintext [OriginChain] <-- one substrate, atomic writes for both shapes | v [API endpoint] <-- ANN search returns ids; hydrate is a parallel batch read ``` One round-trip from the application, no sync layer. ## When Pinecone is the right answer - You already run a primary database that you’re happy with, and you don’t want to migrate it. - You want a dedicated, mature vector engine with very good filter performance at scale. - Your team is comfortable with the dual-write architecture. - Your vector volumes are >100M and growing - Pinecone is more battle-tested at that scale today. ## When OriginChainDB is the right answer - You’re building a new AI feature from scratch and you’d rather not run two databases. - The bugs you spend time on are eventual-consistency bugs between your entity store and your vector store. - You want write atomicity across the entity + the vector + the entity’s secondary indexes. - You want to avoid the ANN re-indexing dance every time entity data changes. ## The atomic-write story The biggest user-facing difference: in OriginChainDB, the entity and its vector are in the same transaction. ```typescript // OriginChain: atomic across shapes await oc.transaction([ { shape: "article", id: "intro-rag", set: { title, body, author } }, { shape: "article-embedding", id: "intro-rag", set: embedding }, ]); // Either both succeed or both fail. The vector is never visible // without the article. ANN search results are guaranteed to point // at hydratable entities. ``` ```typescript // Pinecone: dual-write await db.articles.insert({ id: "intro-rag", title, body, author }); await pinecone.upsert({ id: "intro-rag", values: embedding, metadata: { title } }); // If step 2 fails, the article exists without a vector. Your // ANN search misses recently-inserted content for some window. ``` This isn’t a hypothetical. Every team running this dual-write architecture has at some point debugged a “why isn’t this article showing up in search yet” report. ## The metadata story Pinecone supports filtering on metadata at query time - `filter: {tag: "fintech"}`, etc. Useful, but the metadata schema is independent of your primary database’s schema, and metadata updates are separate operations. OriginChainDB does the same thing structurally - secondary indexes on entity attributes are derived shapes, queryable at any point - but the index is the same data as the entity, not a copy. There’s no second schema to keep in sync. ## Cost comparison | Volume | Pinecone Standard | OriginChainDB entry configuration | | --- | --- | --- | | 1M vectors, low query rate | ~$70/mo | <$20/mo (within free allowance) | | 10M vectors, moderate queries | ~$200-400/mo | $50-150/mo | | 100M vectors, high QPS | $1k+/mo | Talk to us | Pinecone’s pricing scales by both index size and query rate. OriginChainDB’s scales by compute and storage configuration. The comparison depends on workload shape. ## Migration story Migrating from Pinecone to OriginChainDB: 1. Identify the entity. What database holds the canonical record for each Pinecone entry? That database becomes the entity shape in OriginChainDB. 2. Define the vector shape. A sibling shape, e.g. `vec/article/{id}/profile` for an `article` shape. 3. Backfill atomically. For each entity, read it from the source DB, embed (if not already done), and write both shapes in one OriginChainDB transaction. 4. Cut over reads. Replace Pinecone search calls with OriginChainDB ANN search; replace the hydrate-by-id step with a parallel batch read. 5. Decommission the sync layer. The stream/worker that kept Pinecone in sync goes away. ## FAQ ### Is OriginChainDB’s ANN as fast as Pinecone’s? At small-to-medium scale (under 10M vectors, default ANN settings), yes - sub-millisecond p50, ~5ms p99. At 100M+ scale, Pinecone has more tuning headroom today. We’re closing that gap with the optimiser work in step 3 of the [depth-first roadmap](https://originchaindb.com/blogs/depth-first-roadmap-to-1-0). ### Can I run both? Yes - OriginChainDB as the entity + light-vector store, Pinecone as the heavy-ANN engine. We’ve seen this for teams with billion-vector corpuses. ### What ANN algorithm does OriginChainDB use? Currently HNSW with substrate-tuned parameters. Algorithm choice will be configurable per shape post-1.0. ### Does OriginChainDB support filter-then-search like Pinecone? Yes. Combined queries (ANN search + indexed-attribute filter) are first-class - the planner pushes the filter down where it can. ### What about hybrid (sparse + dense) vectors? On the post-1.0 roadmap. Today we support dense vectors only. ## What to read next - [Why we don’t need a separate vector database](https://originchaindb.com/blogs/no-separate-vector-database) - the architectural argument in detail. - [Transactions](https://originchaindb.com/docs/transactions) - what makes the atomic write work. - [One database, every query shape](https://originchaindb.com/blogs/one-database-every-query-shape) - the broader picture. --- # OriginChainDB vs Postgres + pgvector - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/vs-postgres-pgvector Published: 2026-05-05T06:20:43.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal postgrespgvectorcomparisonarchitecturevector-database # OriginChainDB vs Postgres + pgvector Postgres + pgvector suits teams already on Postgres adding vectors as a feature. OriginChainDB suits agent-paced writes and atomic multi-shape transactions. OriginChainDB Team May 5, 2026 About 5 min read Architecture choices / Existing stack ## Match the data model to the team and workload. 1. Existing SQL Schemas and application tooling 2. Vector needs Recall, filtering and scale 3. Query models Text and relationship workloads 4. Migration Compatibility and operational work A database choice includes compatibility, retrieval behavior and the cost of change. Test representative queries instead of choosing from labels alone. TL;DR - Postgres + pgvector is a great fit for teams that already run Postgres and want vectors as a side-quest. OriginChainDB is purpose-built when vectors are first-class and the workload is AI-agent shaped: thousands of small writes per session, atomic multi-shape transactions, sub-millisecond reads. ## TL;DR table | Dimension | Postgres + pgvector | OriginChainDB | | --- | --- | --- | | Primary use case | OLTP with vectors as a feature | AI-native workloads | | Storage substrate | Heap + B-tree | managed K/V | | Atomic vector + row write | Yes (single DB) | Yes (single substrate) | | Schema | DDL, migrations | Key shapes, evolved declaratively | | Throughput ceiling | ~3-30k writes/s on managed | ~1.4M durable writes/s on reference HW | | Vector search at >10M | Slow without IVF tuning | Tuned ANN out of the box | | Operational surface | Tablespaces, vacuum, autovacuum, replication slots | Managed; no operator surface | | When you’d pick it | Stable schemas, mostly humans | Volatile schemas, mostly agents | ## Where Postgres + pgvector wins You already run Postgres. Adding pgvector is a single extension. Your team knows the operational characteristics. Your ORM already speaks SQL. The migration cost is zero, and a Postgres replica is an honest production setup. For a team with Postgres muscle memory and small/medium vector corpuses (under 10M), this is the right answer. Joins matter to your application. Postgres has 30+ years of cost-based-optimiser work behind it. If your access pattern is “vectors and a graph of relational tables hit by `JOIN`s,” nothing else is going to beat it. Your workload is human-paced. Web traffic, dashboards, business apps - Postgres handles thousands of writes per second comfortably and you’ll never feel the ceiling. The per-commit durability cost is invisible at human request rates. ## Where OriginChainDB wins You’re being driven by agents, not humans. Autonomous AI loops emit hundreds of writes and reads per session. With one durable commit at a time, Postgres on managed block storage caps around 3,000 writes/second. OriginChainDB batches many concurrent writers into each durable commit, hitting ~1.4M durable writes/second on reference hardware. That’s two orders of magnitude. Schema is moving. Postgres `ALTER TABLE` on a large table is a real operational event - sometimes online, sometimes not. OriginChainDB’s [key shapes](https://originchaindb.com/docs/schemas) are declarative; adding a field is a config change with no rewrite, no read-only window. Vectors are first-class, not bolted on. pgvector stores vectors in the heap and uses an IVF/HNSW index that you tune yourself. OriginChainDB stores vectors in a dedicated shape co-located with their entity, with the ANN parameters tuned by the substrate. You don’t pick `lists` for IVFFlat or `m` for HNSW - the substrate picks them based on the corpus size and query pattern. That’s less control, but it’s also less wrong by default. You don’t want to operate the database. OriginChainDB is managed-only - no tablespaces, no vacuum, no autovacuum stalls. For a team without dedicated DBAs, the operational delta is large. ## Migration story If you’re migrating from Postgres + pgvector to OriginChainDB, the typical path: 1. Map tables to shapes. Each table becomes one shape (or two if it has secondary indexes worth declaring separately). 2. Extract vectors. Each `vector(N)` column becomes a sibling shape (`vec//{id}/`). Vectors are stored in the same substrate as their entity, so writes stay atomic. 3. Replace `JOIN`s with index lookups. This is the part that takes thinking. OriginChainDB’s access pattern is “key lookup or indexed-attribute lookup or prefix scan.” If your workload depends on multi-table joins with arbitrary predicates, the migration may not pay off - see “where Postgres wins” above. 4. Replace `ALTER TABLE` with shape evolution. Most schema changes become declarative. We’ve seen the migration pay off when the workload was already mostly key-lookup-shaped - the `JOIN`s were doing work that could be modeled as a derived index shape. ## When neither fits If your workload is heavy analytics (long-running aggregations, columnar scans, window functions over billions of rows), neither Postgres nor OriginChainDB is the right primary. You’d want a columnar engine (BigQuery, Snowflake, ClickHouse, etc.) and use either of these as the source-of-truth that feeds it. ## FAQ ### Is OriginChainDB a Postgres clone? No. The substrate is managed key-value, not heap + B-tree. Wire compatibility with Postgres isn’t on the roadmap. The query API is typed but it’s not SQL. ### Can pgvector handle production AI workloads? Yes, up to a point. Teams hit pain past ~10M vectors with high-cardinality filters, or past ~10k writes/sec under sustained agent load. Below that, pgvector is fine. ### Does OriginChainDB support transactions across multiple shapes? Yes. Multi-shape writes commit in a single atomic operation - every shape or none. See [Transactions](https://originchaindb.com/docs/transactions). ### How do I run OriginChainDB locally for dev? You don’t. It’s managed-only. The free database is the dev environment. ~$0.50/day for a sandbox once you outgrow it. ### What about cost? Postgres + pgvector on a managed host: $50-500/month for a small-medium workload, and your bottleneck is IOPS. The OriginChainDB free database costs nothing; paid configurations start at $49/month. At larger scale the comparison flips depending on workload shape - talk to us. ## What to read next - [Why we don’t need a separate vector database](https://originchaindb.com/blogs/no-separate-vector-database) - the architectural argument. - [Schemas](https://originchaindb.com/docs/schemas) - the data model OriginChainDB replaces DDL with. - [Benchmarks](https://originchaindb.com/benchmarks) - the throughput and latency runs we have published. --- # OriginChainDB vs Redis: when each fits - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/vs-redis Published: 2026-05-06T10:26:12.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal rediscomparisonkey-valueperformancearchitecture # OriginChainDB vs Redis: when each fits Redis serves point reads from RAM; its default AOF setting can lose a second of writes. Every OriginChainDB commit is durable before the API returns. OriginChainDB Team May 6, 2026 About 6 min read Architecture choices / Persistence ## Ask what must survive, and how it is read. 1. Read pattern Point access and latency needs 2. Memory Working-set requirements 3. Persistence Acknowledgment and recovery 4. Query models Search and relationship needs Persistence mode, deployment topology and recovery procedures affect the contract. Compare configured behavior, rather than assuming guarantees from a product name. TL;DR - Redis is the gold standard for in-memory KV with sub-millisecond reads. OriginChainDB matches Redis on read latency at small scale and exceeds it on durability semantics, atomic multi-shape writes, and built-in vector search. Redis is the right answer if your data fits in RAM and you accept the persistence story; OriginChainDB is the right answer if you want a primary-store with the same speed feel and AI-native shapes built in. ## TL;DR table | Dimension | Redis | OriginChainDB | | --- | --- | --- | | Latency p50 (point read) | served from RAM | ~250-500 µs | | Persistence | RDB snapshots / AOF; tunable but not the default | durable on every commit | | Atomic multi-key writes | MULTI/EXEC + Lua | Native via shape transactions | | Vector search | RediSearch module (paid tier or self-host) | Built-in shape | | Schema | Schemaless (data structures only) | Declarative key shapes | | Data fits in | RAM (mostly) | Disk + tiered cache | | Operational footprint | You run it, or pay a hosted Redis vendor | Managed only | ## Where Redis wins Pure-RAM speed. Redis serves point reads without touching disk. If your access pattern is “lookup a key, get back a value, repeat at 100k QPS,” Redis is the lowest-latency answer money can buy. Mature data structures. Redis has lists, sets, sorted sets, hyperloglogs, bitmaps, streams, pub/sub, geospatial - battle-tested for 15 years. If your application uses Redis-specific structures (sorted sets for leaderboards, streams for queues), that vocabulary doesn’t translate cleanly to a generic KV substrate. Existing operational muscle. Most engineering teams have a Redis-shaped hole already. Adding another Redis instance is a known cost; switching to a new database is a different conversation. ## Where OriginChainDB wins Durability without ceremony. Redis’s persistence story is “configure RDB snapshots at the right cadence, plus AOF if you really care.” Get it wrong and a power loss costs you whatever’s in the AOF buffer. OriginChainDB writes are durable on commit by definition - there’s no “fast mode” that loses data on power loss because the substrate doesn’t expose one. Vectors without a sidecar. RediSearch (the module that adds vector + full-text + secondary indexes to Redis) is excellent but it’s a separate paid tier on Redis Cloud, or self-hosted complexity if you go open-source. OriginChainDB ships vector search as a first-class shape - no module, no second tier, no second config. Atomic multi-shape writes. Redis MULTI/EXEC gives you atomicity across multiple commands, but they all run on the same instance and you’re explicit about the boundaries. OriginChainDB’s atomic transactions span shapes (e.g. update a user record + their embedding + an index entry) in one atomic operation. The mental model is closer to a real database than to “queue these commands for later.” Schema you can actually evolve. Redis is schemaless in the strictest sense: every reader handles every shape variant. That works at small scale and breaks at large scale when the variants accumulate. OriginChainDB’s [key shapes](https://originchaindb.com/docs/schemas) are declarative and enforced at write time. ## The persistence story, in detail This is where AI-native workloads expose Redis the most. Redis with default persistence: - RDB: snapshot every N seconds. Crash between snapshots = lose those N seconds. - AOF: append every command to a log. Tunable fsync (`always` / `everysec` / `no`). The default is `everysec` - losing up to 1 second of writes on power loss. - Both: more durable but neither is “every commit is durable.” For caching workloads this is fine. For an AI-agent loop that’s writing tool-call traces at 1k/sec, “lose up to 1 second of writes” is hundreds of records you’ll never get back. In OriginChainDB every committed write is durable on disk before the API returns success. There is no “every second” mode; durability is the contract. ## Are we slower than Redis? A bit, at small scale. Redis answers point reads out of RAM; OriginChainDB sits at ~250-500 µs p50 from a co-located client. The difference is the durable write path on the durability side and the storage layer on the read side. For 95% of AI workloads this gap is invisible - your LLM call is going to take 500ms, you spent 5x more on token cost than on the database round-trip, and the 200 µs delta vanishes in the noise. For the 5% where you genuinely need <200 µs (a hot cache in front of a slower primary), Redis is still the right answer. There’s no shame in running both: OriginChainDB as the durable primary, Redis as the read cache. ## Migration story If you’re running Redis as a cache (the most common case), there’s nothing to migrate - keep doing that. OriginChainDB doesn’t try to displace caches. If you’re running Redis as a primary store (which is how you get in trouble with persistence), the migration usually pays off: 1. Map Redis hashes to OriginChainDB shapes. Each Redis hash type → one shape. 2. Map sorted sets / lists / streams as separate shapes with the appropriate access pattern. Sorted sets become an indexed shape with a numeric secondary index. Streams become time-keyed shapes with prefix scans. 3. Write atomicity: replace MULTI/EXEC blocks with multi-shape transactions. The mental model matches. 4. Cut over reads: same client patterns, different SDK. We’ve seen the cleanest wins when the team was using Redis-as-primary with the persistence config that “should be safe enough.” It usually wasn’t, and they didn’t know. ## FAQ ### Is OriginChainDB a Redis replacement? For cache workloads, no - keep using Redis. For Redis-as-primary workloads, often yes - durability + vectors + atomic multi-shape writes are usually the things that pushed you off Redis in the first place. ### Can I use both? Absolutely. OriginChainDB as the durable primary; Redis as a hot read cache for the latency-sensitive 5% of access patterns. ### Does OriginChainDB support Redis data structures (sorted sets, streams)? Not as named primitives. Most can be modelled as shapes - sorted sets as indexed shapes, streams as time-keyed shapes - but you’d write the shape declaration, not run `ZADD`. ### What’s the latency gap going to look like in 12 months? We expect the optimiser work in the depth-first roadmap (step 3) to close the read-side gap meaningfully. Won’t beat pure RAM, but the difference should be small enough that durability + atomic writes win on net. ### Why didn’t you just write a Redis-compatible wire protocol? Wire compatibility constrains the substrate to fit Redis’s data model. Our point is the substrate has a different shape (one store, atomic multi-shape) - wrapping it in RESP would be misleading. ## What to read next - [Schemas](https://originchaindb.com/docs/schemas) - the data model that replaces Redis’s primitive vocabulary. - [Benchmarks](https://originchaindb.com/benchmarks) - the throughput and latency runs we have published. - [Per-key TTL](https://originchaindb.com/blogs/per-key-ttl-agent-memory) - the closest analog to Redis’s EXPIRE. --- # OriginChainDB vs Supabase: when each fits - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/vs-supabase Published: 2026-05-06T10:26:12.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal supabasecomparisonpostgresarchitectureai-native # OriginChainDB vs Supabase: when each fits Supabase bundles Postgres with auth, storage, realtime and edge functions. OriginChainDB is only the database layer, built for AI-shaped data. Where each fits. OriginChainDB Team May 6, 2026 About 5 min read Architecture choices / Product scope ## Separate platform needs from database needs. 1. Application Authentication and service needs 2. Ecosystem Frameworks and existing tooling 3. Database Record and retrieval workloads 4. Ownership What the team will operate Compare the scope of the platform with the scope of your database requirements. The surrounding application services are part of that decision. TL;DR - Supabase bundles Postgres + Auth + Storage + Realtime + Edge Functions into a single managed offering, which is what makes it brilliant for prototyping. OriginChainDB is a focused database substrate purpose-built for AI workloads. If you want a “batteries included” stack, Supabase is the right answer; if you want the database to be excellent at AI-shaped data and you’ll bring your own auth/storage/realtime, OriginChainDB is the right answer. ## What each is for Supabase is a developer platform. Postgres is the core, but the value is the surrounding ecosystem: a real auth service, a file storage layer, realtime subscriptions, edge functions, a generated REST API, and a client SDK that ties everything together. You sign up, get a project URL, and your prototype is live in 30 minutes. OriginChainDB is a database. It does one thing - be the AI-native data substrate - and assumes the rest of your stack (auth, storage, realtime, functions) lives elsewhere. These are different products with different audiences. ## Where Supabase wins Bundle convenience. If you’re building a SaaS prototype and you need auth + DB + a way to upload files + a realtime channel - Supabase gives you that out of the box. Building each piece separately costs days of engineering for a marginal benefit. Postgres compatibility. Anything that talks Postgres talks Supabase. Existing tools, existing libraries, existing intuitions. The ramp is zero. The AI feature lift. Supabase added pgvector + a vector search workflow + AI templates + edge function patterns for OpenAI calls. If your AI feature is “embed this text and store the vector alongside the row,” Supabase covers the path well. Open source. You can self-host Supabase if you want. The “managed” version is a hosting tier, not a wall around the technology. ## Where OriginChainDB wins Vectors as first-class, not as an extension. Supabase ships pgvector; OriginChainDB stores vectors as a [native shape](https://originchaindb.com/docs/schemas) in a managed engine. The performance characteristics are different at scale - see the [Postgres + pgvector comparison](https://originchaindb.com/blogs/vs-postgres-pgvector) - but the mental model is also different: pgvector is a Postgres extension; OriginChainDB’s vector engine reads from the same substrate as the entity write. Throughput shape for agent workloads. Supabase Postgres caps around 3-30k writes/sec on its plans (depending on tier and IOPS). OriginChainDB’s high-throughput write path targets ~1.4M durable writes/sec on reference hardware. The delta matters when an autonomous agent loop is firing hundreds of writes per session at thousands of users. Atomic multi-shape writes by default. In Supabase, atomicity across “the row” and “its vector” is a Postgres transaction - fine. Atomicity across “the row” + “its vector” + “an index entry on a derived attribute” + “an event for downstream subscribers” gets harder to reason about. OriginChainDB’s transaction primitive spans all shapes natively in one atomic write. No bundled auth means you can pick the right one. This sounds like a downside but it’s a feature when you’ve outgrown Supabase Auth (its session model, its rate limits, its specific JWT shape). With OriginChainDB you bring your own - Auth0, WorkOS, Clerk, in-house - and you’re not locked into the bundle’s choices. ## The “batteries included” tax Supabase’s bundle is great until your needs diverge from the bundle’s defaults. Common pain points teams hit at scale: - Auth limitations. Supabase Auth handles signups well but its session/refresh model has constraints that complex apps outgrow. - Realtime fan-out. The realtime layer is fine for low-traffic prototypes; sustained high-throughput broadcast hits limits earlier than dedicated services. - Edge Functions. Convenient but Deno-only and with cold-start characteristics that vary. - Postgres at AI scale. This is the one we’ve already covered: pgvector + agent-shaped writes is a real ceiling. If you’re past these ceilings on multiple dimensions, the bundle stops being a discount and starts being a constraint. OriginChainDB is opinionated about being the database layer specifically - pick the right tool for the rest, rather than assuming one stack vendor is best at everything. ## Migration story The realistic migration from Supabase to OriginChainDB isn’t “swap the whole stack.” It’s “swap the data layer for AI-shaped workloads, keep the rest of Supabase for everything else.” Pattern we’ve seen: 1. Keep Supabase Auth + Storage + Realtime. 2. Move the AI-feature data (entity + vectors + signal) to OriginChainDB. 3. Connect them via your application code: Supabase Auth provides the session, your code reads/writes both Supabase Postgres (for human-CRUD data) and OriginChainDB (for AI data). This is the “two databases” architecture you’d otherwise have if you used Postgres + Pinecone - except OriginChainDB replaces both halves of that pair and gives you the atomic-write story. ## When to stay on Supabase - The AI feature is small relative to the rest of the app. Don’t add a database for one feature. - Your write rate is well under 3k/sec sustained. Postgres handles that fine. - You’re still in prototype phase and the velocity from the bundle outweighs the ceiling concerns. - The bundled features (auth, realtime, storage) are doing real work for you. ## FAQ ### Is OriginChainDB a Supabase competitor? For the data layer specifically, yes. For the bundle (auth + storage + realtime + edge functions), no - we’re not building those. ### Can I run OriginChainDB alongside Supabase? Yes. The most common pattern is Supabase for the human-app data + auth + realtime; OriginChainDB for AI-feature data + vectors. Your application talks to both. ### Does OriginChainDB have an auth layer? Not as a product surface. The admin panel has its own auth (per-user magic links + role-based access), but customer applications bring their own. ### What’s the equivalent of Supabase’s “instant API”? OriginChainDB’s typed SDK + OpenAPI spec covers the same role. Different shape - typed methods rather than auto-generated REST - but the “no API server to write” story is the same. ### Will OriginChainDB ever bundle auth/storage/realtime? No, by design. Depth-first to 1.0 means doing the database well; bundle features are post-1.0 if at all. ## What to read next - [vs Postgres + pgvector](https://originchaindb.com/blogs/vs-postgres-pgvector) - Supabase’s database engine on its own merits. - [Schemas](https://originchaindb.com/docs/schemas) - the data model. - [Why we don’t need a separate vector database](https://originchaindb.com/blogs/no-separate-vector-database) - the architectural reason for picking OriginChainDB over Postgres + pgvector for AI workloads. --- # Window functions, correlated subqueries - OriginChainDB Blog Canonical source: https://originchaindb.com/blogs/window-functions-correlated-subqueries-590-tests Published: 2026-06-06T10:00:00.000Z (historical article; not a current availability promise) Sitemap last modified: 2026-09-13T14:38:26.000Z [All insights](https://originchaindb.com/blogs) OriginChainDB journal sqlwindow-functionscorrelated-subqueriesengineeringplanner # Window functions, correlated subqueries OriginChainDB's SQL surface now has ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD and aggregate OVER, plus correlated EXISTS, IN and scalar subqueries. OriginChainDB engineering Jun 6, 2026 About 7 min read Query engineering / SQL ## Keep the row, add context from its group. 1. Partition Choose the row's group 2. Order Define a sequence in that group 3. Compute Rank, compare or aggregate 4. Return Keep row-level detail Window functions add calculations without collapsing every group into a single row. Correlated subqueries express a different relationship to each outer row. TL;DR — OriginChainDB’s SQL surface just grew window functions and correlated subqueries: ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD and SUM/AVG/COUNT/MIN/MAX OVER, plus correlated EXISTS, IN and scalar subqueries. Both are first-class on the planner. Running aggregates are limited to the cumulative frame — an explicit `ROWS BETWEEN` is refused with a hint and queued for v2. Two SQL surfaces landed in the engine this cycle. They’re the kind of features that sound boring on a roadmap line and matter the moment a customer pastes a real analytics query into our SQL endpoint. Around 590 tests across the SQL translator, the query executor, and the HTTP layer now run on every push — the surface is large enough that we stopped trusting cursory review and started leaning on the suite. ## What works in v1 Six window function families. Every one of them runs against a plain single-table SELECT, with PARTITION BY and ORDER BY inside the OVER clause: - `ROW_NUMBER() OVER (...)` — per-partition counter, deterministic by ORDER BY - `RANK() OVER (...)` — competition ranking (gaps on ties) - `DENSE_RANK() OVER (...)` — dense ranking (no gaps) - `LAG(col [, offset [, default]]) OVER (...)` — peek backwards in the partition - `LEAD(col [, offset [, default]]) OVER (...)` — peek forwards - `SUM / AVG / COUNT / MIN / MAX (col) OVER (...)` — running aggregates over the partition Correlated subqueries cover the three shapes that show up in real customer SQL: - `EXISTS (SELECT ... WHERE inner.k = outer.k)` — including the `NOT EXISTS` negation, rewritten to a semi-join - `WHERE col IN (SELECT ... WHERE inner.k = outer.k)` — correlated IN - Scalar `(SELECT col FROM ... WHERE inner.k = outer.k) = value` — correlated scalar predicate The translator detects correlation by looking at which columns the inner SELECT references. Uncorrelated subqueries take a fast path that materializes once; correlated subqueries fall through to a per-outer-row execution that the planner stitches into the predicate tree. ## A worked example The canonical “most expensive order per customer” query, the way you’d write it against any real OLTP database: ```sql SELECT customer_id, order_id, total_cents FROM ( SELECT customer_id, order_id, total_cents, ROW_NUMBER() OVER ( PARTITION BY customer_id ORDER BY total_cents DESC, order_id ) AS rn FROM orders ) WHERE rn = 1; ``` This runs against OriginChainDB today, against our SQL endpoint, with the same shape you’d send to Postgres. The planner turns the inner query into: ```plaintext Scan(orders) -> Window(ROW_NUMBER, partition_by=[customer_id], order_by=[total_cents DESC, order_id]) -> ProjectAliased(customer_id, order_id, total_cents, rn=row_number) ``` The Window operator does a per-partition counter walk in the executor; the outer filter `WHERE rn = 1` is pushed onto the projected stream. No special-cased planner hack, no recipe match — `ROW_NUMBER OVER` is just a Plan node now. The same shape gives you “yesterday’s value alongside today’s” with LAG: ```sql SELECT symbol, ts, price, LAG(price, 1, 0.0) OVER ( PARTITION BY symbol ORDER BY ts ) AS prev_price FROM ticks; ``` And running totals with SUM OVER: ```sql SELECT user_id, event_ts, amount_cents, SUM(amount_cents) OVER ( PARTITION BY user_id ORDER BY event_ts ) AS lifetime_spend FROM payments; ``` Both work. Both have tests. Both are the queries we expect customers to actually run. ## The frame trade-off Every running aggregate in v1 runs at the cumulative frame — implicit `RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW`. If you write an explicit `ROWS BETWEEN 2 PRECEDING AND CURRENT ROW`, we refuse with a clear hint rather than silently ignore the clause. This is deliberate. The implementation cost of explicit frames is real — the executor needs a sliding-window aggregator, with the per-frame add/evict bookkeeping that comes with it. Doing that well is a quarter of engineering on its own. Shipping a partial frame implementation that quietly produces wrong numbers for `ROWS BETWEEN n PRECEDING` would be worse than refusing. So v1 ships the cumulative frame, and v1 refuses everything else with a message that names the missing surface and tells the caller to wait for v2. The two-line failure on an unsupported frame is a feature, not a regression. What this means in practice: - Running totals, running maxes, running counts: work. - Moving averages over a sliding window: refused, v2. - Three-day moving average over price ticks: refused, v2. If your analytics workload is mostly cumulative aggregates and ranking, v1 covers you. If it’s heavy on sliding-window analytics, v1 doesn’t. ## What `EXISTS` looks like under the planner The fun part of correlated subqueries is what they compile to. `EXISTS (SELECT 1 FROM orders WHERE orders.customer_id = customers.id)` is, structurally, a semi-join: “keep the customer row if at least one matching order row exists.” The translator detects the correlation by walking the inner SELECT’s predicate tree, finds the `inner.customer_id = outer.id` binding, and emits a Semi-Join plan node rather than a per-row subquery execution. `NOT EXISTS` becomes an Anti-Join. The same detection logic runs, and the planner inverts. This matters because the Semi/Anti-Join path is orders of magnitude cheaper than naive per-outer-row execution. A customer table with 100k rows joined against orders with 1M rows runs as one hash-join probe, not 100k separate subquery executions. The fallback path — when the correlation is too gnarly for the semi-join rewrite — does run the subquery once per outer row. We name this in the planner output. Subqueries in WHERE (IN / EXISTS / scalar) return 400 in PREVIEW - write the query as an explicit JOIN. ## What’s queued for v2 Five gaps we know about. We’re naming them so a reader evaluating the surface doesn’t have to guess: - Explicit frames — `ROWS BETWEEN n PRECEDING AND CURRENT ROW` and friends. Required for moving averages. - `FILTER (WHERE ...)` — the SQL standard filter clause on window aggregates. Currently refused with a hint. - Named windows — `OVER win_name` referencing a `WINDOW win_name AS (...)` clause. Currently refused; inline the spec into each OVER call. - `FIRST_VALUE` / `LAST_VALUE` / `NTH_VALUE` / `NTILE` / `PERCENT_RANK` / `CUME_DIST` — the analytic functions we didn’t ship in v1. Tracked. - Window + GROUP BY / JOIN in the same SELECT — the planner refuses this combination today. The SQL-standard answer is “window runs after aggregation” — for now, do the aggregate or join in a subquery and apply the window to the result stream. Every refusal in the translator carries a hint that names what to do instead. We’d rather a customer see a one-line “use a subquery first” message than ship a wrong answer. ## The test count, honestly Around 590 passing tests across the SQL-relevant components. The split: - SQL translator: 269 tests — the translator from SQL AST to plan tree. - Query executor: 199 tests — the executor that runs the plan tree. - HTTP surface: 123 tests — the layer that turns `POST /v1/tenants/{tenant}/sql` into a typed response. This is the number we trust on a green CI run. We don’t claim 100% coverage — there is no static analysis behind it — but every window-function and correlated-subquery code path has at least one positive test, one refusal test, and one round-trip test against a real executor. The refusal tests matter more than the positive ones: they’re what guarantee we don’t silently regress into a wrong answer when somebody adds a new SQL feature. ## When to use this surface The window functions make OriginChainDB’s SQL surface usable for the analytics queries that show up in real customer workloads: - Personalization features — “find each user’s most-recent N items” or “rank items per category by score.” Window functions are the right shape. - Cohort and ranking dashboards — running aggregates over a partition. - Existence and uniqueness checks against a derived set — `WHERE EXISTS (subquery)`. The correlated path is the natural shape and the planner now handles it. If your workload is heavy sliding-window analytics, wait for v2. If it’s ranking + correlated existence, ship today. ## Try it Provision a tenant from the [quickstart](https://originchaindb.com/docs/quickstart). Open the SQL endpoint. Paste a real query with `ROW_NUMBER OVER (PARTITION BY ...)` or `WHERE EXISTS (SELECT ...)`. If you hit a refusal, the hint tells you what to do instead — and the v2 list above tells you whether the surface is queued. The full supported surface is in the [SQL reference](https://originchaindb.com/docs/sql). --- # Careers | OriginChainDB Canonical source: https://originchaindb.com/careers Sitemap last modified: 2026-09-27T14:44:27.000Z Company / Careers # Build database infrastructure. We’re building a distributed multimodal database. The work spans the engine, the infrastructure and the tools developers use every day. 1. The engine SQL, vectors, full-text and graph relationships. 2. The infrastructure Distributed systems, operations and security. 3. The experience Tools that help developers understand their data. Three parts of the product we’re building Current opportunities ## No open roles right now. Interested in future opportunities? Tell us what you build and share a link to your work. We’ll review expressions of interest when a relevant role opens. [Introduce yourself](https://originchaindb.com/contact?topic=careers) Get to know the work ## Start with the details. ### Understand the architecture Follow how data is stored, queried and connected across models. [Inside the database ↗](https://originchaindb.com/architecture) ### Try the developer experience Explore the documentation and a sample workspace before starting a conversation. [Explore the demo workspace ↗](https://demo-console.originchain.ai/) ### Only confirmed openings. We publish roles here when they are ready. An expression of interest is not an application to an advertised vacancy. Read our [privacy policy](https://originchaindb.com/privacy) before sharing personal information. --- # OriginChainDB release notes Canonical source: https://originchaindb.com/changelog Published: 2026-07-06 (historical article; not a current availability promise) Article modified: 2026-07-06 Release dates: 2026-07-06, 2026-06-08, 2026-05-03, 2026-05-02, 2026-05-01, 2026-04-30, 2026-04-23, 2026-04-22, 2026-03-20, 2026-02-12, 2026-01-20, 2025-12-18. Release notes describe their dated versions, not the current deployment. Sitemap last modified: 2026-09-23T08:33:23.000Z [Resources](https://originchaindb.com/resources) The release archive # What changed. Version by version. Published features, fixes, and technical notes—kept with the release they describe. 13 published entriesDecember 2025 — July 2026[Read the newest entry](https://originchaindb.com/changelog#v1-3-2026-07-06) These notes describe the listed releases. They do not identify the version currently running on your instance. A map of the archive ## Follow the release history. Published version and date sequenceOldest to newest · select a version to read its notes 1. ### Dec 2025 [v0.518](https://originchaindb.com/changelog#v0-5-2025-12-18) 2. ### Jan 2026 [v0.620](https://originchaindb.com/changelog#v0-6-2026-01-20) 3. ### Feb 2026 [v0.712](https://originchaindb.com/changelog#v0-7-2026-02-12) 4. ### Mar 2026 [v0.820](https://originchaindb.com/changelog#v0-8-2026-03-20) 5. ### Apr 2026 [v0.1022](https://originchaindb.com/changelog#v0-10-2026-04-22)[v0.1123](https://originchaindb.com/changelog#v0-11-2026-04-23)[v0.1230](https://originchaindb.com/changelog#v0-12-2026-04-30) 6. ### May 2026 [v1.001](https://originchaindb.com/changelog#v1-0-2026-05-01)[v1.002](https://originchaindb.com/changelog#v1-0-2026-05-02)[v1.0.102](https://originchaindb.com/changelog#v1-0-1-2026-05-02)[v1.103](https://originchaindb.com/changelog#v1-1-2026-05-03) 7. ### Jun 2026 [v1.208](https://originchaindb.com/changelog#v1-2-2026-06-08) 8. ### Jul 2026 [v1.306](https://originchaindb.com/changelog#v1-3-2026-07-06) About the tags feat Features perf Performance fix Fixes docs Documentation breaking Breaking changes Newest first ## v1.3 "Scale out" 2026-07-06 Multi-node configurations reach GA - write throughput that scales with node count behind one endpoint, transactions that span nodes and commit atomically, point-in-time recovery to the roll-up window, and rate budgets that grow with your configuration. Release details5 notes - feat Multi-node configurations are generally available. Pick a node count when you build your configuration: write throughput scales with nodes (~2x measured at 2 nodes), data distributes across nodes automatically table by table, and your application keeps talking to a single endpoint with a single key - no client-side routing, no code changes. Each node can optionally carry its own synchronized standby through the same resilience configurations as single-node. - feat Atomic cross-node transactions. BEGIN / COMMIT / ROLLBACK over the SQL endpoint now spans tables on different nodes with the same all-or-nothing guarantee as a single node: every write in the transaction lands, or none do. ROLLBACK is idempotent and always safe to call; a lost transaction surfaces as a clean 409 with zero partial writes. See the new Transactions reference in the docs. - feat Point-in-time recovery: restore an instance to a roll-up boundary, with integrity verification on the restored data before it serves traffic. Driven from the console or the API. The archive window that bounds how close to a chosen moment you can land is a server-side setting, not a measured recovery point. - feat Rate budgets now scale with node count. Per-key budgets (requests/sec, bytes/sec, natural-language queries/sec, concurrent queries) multiply by the number of nodes in your configuration, and the full aggregate budget is delivered through the single endpoint. 429 responses now carry an X-OC-Limit-Hit header naming exactly which budget you exhausted. - docs New docs pages: Transactions (lifecycle, error handling, the retry pattern) and Multi-node configurations (what you get, choosing a node count, growing later). The rate-limits reference is rewritten around the four per-key budgets - and why batch endpoints (~18x measured throughput vs single-row calls) are the first lever to pull. [Link to this release](https://originchaindb.com/changelog#v1-3-2026-07-06) ## v1.2 "Scale ceiling" 2026-06-08 Vectors scale past 10M per tenant, graphs learn embeddings, full-text reduces to canonical lemmas, and multi-writer cluster replication stays on the roadmap - active-passive remains the only replication mode any shipped engine serves. Release details12 notes - feat IVF and IVF-PQ vector indexes land end-to-end. IVF partitions vectors by inverted-file cells for 10M+ scale per tenant; IVF-PQ adds residualised product quantization per cell (the Jegou 2011 path) for 64× memory savings at D=128 and 768× at D=1536 OpenAI ada-002 dims. The CI acceptance floor is recall@10 >= 0.75 at nprobe=4 on a small planted-cluster corpus. 1M IVF bulk-load in 80 s. - feat Binary quantization (32× memory savings) and PQ quantization (64× at the headline config) attach to HNSW or IVF as separate index kinds. Pick the recall/memory tradeoff per-collection without rewriting your client code. - feat GraphSAGE attribute-aware node embeddings ship with three aggregator choices: Mean (fastest), MaxPool (best representational capacity per parameter), and LSTM (sequence-aware, deterministic per-(node, layer) shuffle so retrains converge to the same vector). Pairs with Node2Vec persistence for downstream similarity search. - feat Full-text lemmatization in 9 languages — English, Spanish, French, German, Italian, Portuguese, Russian, Dutch, Swedish — reduces inflected forms to a dictionary-backed canonical lemma. Higher precision than Snowball stemming on natural-language fields. Stemming for 18 languages remains available as a separate analyzer step. ICU + geo tokenizer ship alongside for mixed-script and location-aware queries. - feat Cypher v3 closes: CALL subqueries, nested FOREACH, list comprehensions (row-time), MERGE, DELETE, CREATE relation, and multi-hop undirected matches. The DELETE / CREATE path lands first; MERGE follows the same idempotent contract you'd expect from a Cypher implementation. - feat Materialized views with on-demand refresh: install_materialized_view computes the initial materialization, refresh_materialized_view re-runs the definition with an atomic overwrite. POST install / POST :name/refresh / GET :name. Incremental refresh stays on the roadmap. - feat Foreign keys and CHECK constraints (Phase B). FK on-delete supports NoAction, Restrict, and SetNull at write-time; Cascade and SetDefault are deferred. CHECK expression language: literals, =/!=//>=, AND/OR/NOT, IS [NOT] NULL, IN — with 3-valued logic. - feat JOIN cap raised from 5 to 32 tables. Left-deep planner; analytics rollups across a long dimension chain no longer hit the ceiling. - feat Multi-writer (quorum) cluster replication is roadmap, not a shipped mode, and has no scheduled date. No shipped engine commits a write through consensus: a request to run a database in quorum mode is refused at config-install time rather than accepted and quietly degraded, and the consensus machinery runs only on engines explicitly configured for engineering drills. Active-passive replication is the supported mode and the default for every tenant. - feat Serializable MVCC Phase B (preview): per-key version chains with !ts prefix-scan-from-newest give O(1) conflict detection at ~28-byte overhead per version. Range-serializability via gap locks is Phase C work. - feat PostgreSQL ingest connector v1 (POST /v1/.../ingest/postgres/sync) pulls rows from an existing Postgres source into your OriginChainDB tenant — no separate ETL service needed. - feat Python SDK adds typed namespaces: oc.sql, oc.vector, oc.fts, oc.graph. IDE autocomplete works against the shape you're actually using. [Link to this release](https://originchaindb.com/changelog#v1-2-2026-06-08) ## v1.1 "Hardening" 2026-05-03 Transactional email, crash-injection in CI, shape-level diagnostics, and a fistful of dashboard fixes. Release details16 notes - feat Forgot- and reset-password flow shipped end-to-end. Request a reset link from /forgot, set a new password from /reset-password - single-use, time-limited tokens delivered through our managed transactional email service. - feat Magic-link login and password-reset emails now go through our managed transactional email service. Delivery is signed, DKIM-authenticated, and falls back to a journal log if the provider has a hiccup so you never lose access. - feat Crash-injection harness lives in CI. Every commit crashes the engine at four critical durability boundaries twenty times each - recovered state must be a prefix of a known-good model. 4×20 = 80 green on every push. - feat Optimiser landed on the hot path: cost-based shape routing now picks the cheapest plan from a small fixed menu instead of always taking the first match. No tuning knobs to learn. - feat Synchronous-replica catch-up now uses an explicit sync state-machine - followers report replication delay in real time and the primary refuses to ack writes until the standby is caught up. - feat Configuration-upgrade flow in the dashboard: change compute configuration from Single-zone → HA → HA+ without recreating the instance. Rolling resize, no downtime, prorated on the next bill. - fix Billing math: monthly spend now includes only running instances. Stopped instances no longer accrue compute charges in the current-cycle estimate (storage still does, since data still lives on disk). - fix Admin login: the auth-gate IIFE no longer fires on / or /magic-link, so the reload loop reported on a few accounts is structurally impossible. Login form is always interactive, even before boot-status comes back. - fix Dashboard /app/backups: hoisted the tenant-ownership check out of an N+1 loop and translated AbortError into a friendly retry banner. List-backups no longer SIGNAL_ABORTED on instances with hundreds of backups. - fix Console snapshot list now caches results for 30 s and the SDK times out at 60 s - slow snapshot directories no longer hang the page. - fix Public docs: scrubbed a leaked tenant ULID and replaced with the acme.* placeholder used elsewhere. - fix Crash-replay test corpus expanded - every persisted shape now has a deterministic replay test. - perf Per-repo CI: each crate, each SDK, and the frontend run in their own pipeline. Builds that don't touch a target skip its tests entirely - typical PR feedback dropped from ~9 min to ~2 min. - docs New blog post: "Crash at every durability boundary" - how the new crash-injection harness works and why it caught two real bugs the day it landed. - docs New blog post: "One commit, every shape" - how multi-shape writes commit atomically as a single durable unit. - docs New blog post: a deep-dive on how the cost-based optimiser picks shape routes. [Link to this release](https://originchaindb.com/changelog#v1-1-2026-05-03) ## v1.0.1 "Outreach" 2026-05-02 Marketing surface, MCP server, and an OpenAPI 3.1 spec - built so AI agents and humans can both read OriginChainDB. Release details11 notes - feat @originchain/mcp-server published - a Model Context Protocol server that exposes /rows, /query, /ask, /watch, and the add-on shapes to any MCP-aware client (Claude Desktop, Cursor, etc.). Drop in your bearer, point at your endpoint, you're live. - feat OpenAPI 3.1 spec published at /openapi.json - every public HTTP endpoint, every schema, every error code. Generate clients in any language, or paste it straight into your editor's HTTP plugin. - feat Eight head-to-head comparison pages: OriginChainDB vs Postgres, MongoDB, Neon, Supabase, Pinecone, Qdrant, Weaviate, and Milvus. Honest about where each tool wins and where ours does. - feat Site-wide search via Pagefind. Hit the search input on any page, get instant results across docs, blogs, and marketing pages - fully static, no backend, works offline once cached. - feat Real /status page wired to the same metrics the on-call dashboard reads. Per-region uptime, 30-day incident history, current SLO burn. - feat Schema delete in the dashboard: drop a key shape from the console with a typed-confirmation guard. Backed by new DELETE endpoints (DELETE /v1/schemas/:name with cascade flag). - feat Per-page Open Graph images: every marketing page renders a tailored OG card with its own title and category, so links look right on Slack, Twitter, LinkedIn, and Discord. - feat Mobile hamburger navigation across the marketing site, with dark theme as the default. Light theme toggle still available. - feat IndexNow ping on every deploy - search engines hear about new docs and blog posts within seconds of merge. - feat Three new long-form engineering blog posts on storage internals, replication, and the new add-on architecture. - fix Go SDK: corrected the module path to github.com/originchain-ai/originchain-go. `go get` and `go mod tidy` now resolve correctly. [Link to this release](https://originchaindb.com/changelog#v1-0-1-2026-05-02) ## v1.0 "Add-ons" 2026-05-02 Six new paid add-ons attach to any configuration - opt in only what your workload needs. A seventh, the multi-writer cluster, was announced here but has not shipped. The prices below are the ones announced that day and are no longer the offer: SQL Pro was folded into every paid configuration on 2026-06-03, and Vector Search, Full-Text Pro and Graph followed on 2026-09-02 - all four are now included at no extra charge. MVCC Transactions and Intra-Segment PITR are still priced add-ons. A configuration is priced as a whole today - see [pricing](https://originchaindb.com/pricing) for what applies now. Release details7 notes - feat SQL Pro add-on ($49/mo): GROUP BY with COUNT/SUM/AVG/MIN/MAX, INNER + LEFT/RIGHT/FULL OUTER joins, HAVING filters, and chained 3+ table joins (left-deep, up to five tables). - feat Vector Search add-on ($79/mo + $0.0002 per topk): HNSW with cosine, dot, L2 metrics and tunable speed/recall. Default high_recall mode hits recall@10 = 0.96 at 100k vectors with p99 109 ms; fast mode runs p99 37 ms at recall 0.69. Filtered topk via metadata equality, f32 SIMD distance kernels, deserialized graph cache. - feat Full-Text Pro add-on ($49/mo): Lucene-default BM25 (k1=1.2, b=0.75), phrase queries via position-list intersection, UAX #29 Unicode tokenizer, and Snowball stemming for 18 languages. - feat Graph add-on ($59/mo): forward and reverse one-hop neighbors (correct on self-relations), BFS up to a configurable max depth, path reachability, and weighted shortest path (Dijkstra) with caller-supplied weight functions. - feat Transactions add-on ($99/mo): multi-row snapshot-isolation transactions with optimistic conflict detection - begin / get / put / delete / commit / abort across multiple rows. - feat Intra-Segment PITR add-on ($149/mo): tighter point-in-time recovery via a continuous change stream with embedded microsecond timestamps. Preview - the archive window that bounds restore granularity is a server-side setting, not a measured recovery point. - feat Multi-Writer Cluster add-on (Enterprise): announced here as a multi-node write cluster with TLS transport and durable persistence. It has not shipped. Quorum replication remains roadmap with no scheduled date, a request to run a database in that mode is refused at config-install time, and every database in service runs the active-passive path. [Link to this release](https://originchaindb.com/changelog#v1-0-2026-05-02) ## v1.0 "Live deploy" 2026-05-01 OriginChainDB reaches general availability - single-tenant, region-isolated, sub-100 ms in-region p99. Release details5 notes - feat Single-tenant region-isolated compute live in Asia Pacific (Mumbai). Each instance gets a dedicated instance, durable transaction logging, and an HTTPS endpoint provisioned in about two minutes. - feat Standby replication on every paid configuration. A write is acknowledged once it is flushed to durable storage on the primary; frames stream to the standby continuously, asynchronously. - feat Three compute configurations (Single-zone, HA, HA+) and four storage configurations (20 GB to 2 TB). Resize compute and storage independently, hot, with no downtime. - feat Console-driven bearer rotation with a 60-second grace window for rolling deploys. - perf p99 read latency under 8 ms in-region on HA and HA+; cached /ask responses return in under 50 ms. [Link to this release](https://originchaindb.com/changelog#v1-0-2026-05-01) ## v0.12 "Failover" 2026-04-30 Promoted-follower failover verified end-to-end on live infrastructure. Release details4 notes - feat Snapshot-bootstrap: a new follower can join a running primary without operator intervention, receiving a consistent snapshot before tailing live frames. - feat Failover preserves full state - promoted follower serves reads and writes immediately on takeover. - perf Promotion is operator-triggered by default. An opt-in automatic promotion path also exists: it is off by default, a standby takes the writer role only once the writer lease has expired, an external restart hook completes the switch, and a standby that is not fully caught up refuses to promote. - perf Active-passive replication remains the production path. Promotion is fenced by a single-primary claim; because replication is asynchronous, an abrupt primary loss can still cost the most recent acknowledged writes. [Link to this release](https://originchaindb.com/changelog#v0-12-2026-04-30) ## v0.11 "Initial GA" 2026-04-23 Engine, managed console, and Razorpay billing go live. Release details4 notes - feat Managed console: signup, region selection, configuration selection, instance provisioning, metrics and audit log access. - feat Razorpay billing integration: USD pricing, monthly and annual cycles, prorated mid-cycle changes, hosted card capture at instance create. - feat Encrypted nightly backups with daily integrity checks; continuous-archive point-in-time recovery on every paid configuration. - feat Python SDK - fully typed, sync and async, with helpers for /rows, /query, /ask, /watch, /sql, /vector, /fts, /graph. [Link to this release](https://originchaindb.com/changelog#v0-11-2026-04-23) ## v0.10 "Compile" 2026-04-22 Natural-language compile path runs fully in-region with no API key in the customer's path. Release details4 notes - feat Every /ask call compiles in the same region as the tenant's instance. Credentials are scoped per instance and handled for you - no LLM API keys anywhere in the customer's stack. - feat Schema prompt caching: repeated /ask calls with the same catalog pay the cached-read rate (~90% discount on repeated catalog tokens). - feat Per-instance scope locked to a single model identifier - no other foundation models, no cross-tenant access. - perf Shared runtime + per-region inference client cache; SDK startup cost is paid once per process, not per /ask. [Link to this release](https://originchaindb.com/changelog#v0-10-2026-04-22) ## v0.8 "Production" 2026-03-20 Hot-path concurrency, atomic backups, chaos and load harnesses. Release details5 notes - feat MVCC-lite on hot-path writes - conflicting updates serialize per row without blocking reads. - feat oc backup create / oc backup restore - atomic at the filesystem level, with daily integrity checks. - feat Chaos and load harnesses run on every build (process kill, disk full, interrupted-write recovery; 10k rps steady-state). - feat HTTP body limit enforced at 8 MiB; rate limiter is a real token bucket with Retry-After headers. - perf Concurrent-read path reworked to remove a hot-index bottleneck - p50 reads down ~18%. [Link to this release](https://originchaindb.com/changelog#v0-8-2026-03-20) ## v0.7 "Reactive" 2026-02-12 Live views over Server-Sent Events; durable plan cache. Release details3 notes - feat New /v1/watch endpoint - Server-Sent Events stream of updates for a live expression. Same protocol across every region and configuration. - feat Plan cache spills to disk; cold start reloads hot plans without re-invoking the LLM. - perf Catalog lookup is lock-free; consult cost dropped to ~12 µs. [Link to this release](https://originchaindb.com/changelog#v0-7-2026-02-12) ## v0.6 "Consolidation" 2026-01-20 Single managed engine replaces the multi-engine split. Release details3 notes - breaking Storage consolidated onto a single managed engine - one recovery path, one backup path, one mental model. Earlier multi-engine split removed. - breaking Parquet, Arrow, DataFusion, and KuzuDB dependencies dropped. - feat Plans compile down to index, ref-walk, and range-scan primitives on the unified store. Cost-based routing retired - no routing to do. [Link to this release](https://originchaindb.com/changelog#v0-6-2026-01-20) ## v0.5 "First release" 2025-12-18 Initial public release. Release details2 notes - feat HTTP ingress, row-keyed store, and the natural-language to query-plan compiler land for the first time. - docs Specification document published; HTTP API documented. [Link to this release](https://originchaindb.com/changelog#v0-5-2025-12-18) Keep up with the changes ## Get release notes in your inbox. Request release updates by email, or use the docs to explore a feature. [Subscribe by email](mailto:support@originchain.ai?subject=subscribe%20to%20release%20notes)[Open the documentation](https://originchaindb.com/docs) --- # About OriginChainDB | OriginChainDB Canonical source: https://originchaindb.com/company Sitemap last modified: 2026-09-24T06:54:16.000Z The company behind OriginChainDB # Built for the ways your data connects. We're building a distributed multimodal database for records, semantic search, relationships and text. One product for developers working across all four. [Explore the architecture](https://originchaindb.com/architecture)[Talk with the team](https://originchaindb.com/contact) Built bySilicoyn Technologies Pvt Ltd One product, four models ## A common home for different kinds of data. The project began at the end of 2024 with a focus on bringing records, retrieval and relationships into one database product. The database engine is built in Rust. What we buildFour data models [SQL Structured records Filter, join and aggregate](https://originchaindb.com/docs/sql/quickstart)[Vector Supplied embeddings Retrieve similar items](https://originchaindb.com/docs/vector/quickstart)[Graph Named relationships Follow connected records](https://originchaindb.com/docs/graph/quickstart)[Full-text Indexed documents Find words and phrases](https://originchaindb.com/docs/fts) OriginChainDBDatabase engine APIs Console One team, building the whole product.From database internals to the developer experience. Four ways to work with data. Your application supplies embeddings, updates each index and connects results by record ID. Natural-language questions use the optional `/ask` query interface. [See how questions become queries](https://originchaindb.com/docs/ask) How we build ## Make the useful parts understandable. Our focus is the database and the experience of building with it. 01 ### Start with working interfaces. SQL and HTTP examples show what to send, what comes back and how to connect the results to your application. [Run the quickstart](https://originchaindb.com/docs/quickstart) 02 ### Keep the system explainable. Separate the write path, retrieval and deployment choices so teams can reason about their own workloads. [Understand the design](https://originchaindb.com/architecture) 03 ### Evaluate with your data. Use your queries, corpus and deployment requirements to decide whether OriginChainDB fits. [Plan a technical evaluation](https://originchaindb.com/contact) People ## From technical questions to customer conversations. [Image: Zaheer Kazi] Co-Founder & Chief Commercial Officer ### Zaheer Kazi Zaheer leads OriginChainDB's commercial operations, revenue and customer success, bringing more than three decades of enterprise technology experience. His career includes roles at Hewlett-Packard, DXC Technology, Sun Microsystems, Lucent Technologies, Azentio Software and Syngene International. [Connect on LinkedIn](https://www.linkedin.com/in/zaheerkazir/) Company details ## The people to reach. The company to work with. Legal entity ### Silicoyn Technologies Pvt Ltd The operator of OriginChainDB and the counterparty on our agreements. Sales [sales@originchain.ai](mailto:sales@originchain.ai) Support [support@originchain.ai](mailto:support@originchain.ai) Security [security@originchain.ai](mailto:security@originchain.ai) Memberships Member of the NVIDIA Inception program and the AI Council of India (IAMAI). [NVIDIA InceptionProgram member](https://www.nvidia.com/en-us/startups/)[AI Council of IndiaInternet & Mobile Association of India](https://www.iamai.in/) Build with us ## Start with a real use case. Explore the examples, try a database, or bring your requirements to the team. [Start free](https://app.originchaindb.com/signup)[Read the quickstart](https://originchaindb.com/docs/quickstart) --- # Compare database approaches | OriginChainDB Canonical source: https://originchaindb.com/compare Sitemap last modified: 2026-09-23T08:33:23.000Z compare # Compare database approaches Compare data models, query capabilities and operational trade-offs to decide where OriginChainDB fits your application. the frame ## One question decides it: how many systems does a write touch? If your application is a relational app that occasionally needs an embedding, Postgres with pgvector is the right answer. If it is an AI application whose rows, vectors, full-text postings and graph edges must agree on every write, a distributed multimodal engine with one atomic commit removes the sync job — and the second bill — between them. relational primaries [OriginChainDB vs Postgres + pgvector The default primary. Where an extension is enough, and where it stops being one database. Read the comparison →](https://originchaindb.com/vs/postgres)[OriginChainDB vs Neon Serverless Postgres with branching, compared on AI workload shape. Read the comparison →](https://originchaindb.com/vs/neon)[OriginChainDB vs Supabase Postgres plus a platform. Where the platform helps and where the vector path is still pgvector. Read the comparison →](https://originchaindb.com/vs/supabase) dedicated vector databases [OriginChainDB vs Pinecone Managed vector search versus vectors that commit with the rows they describe. Read the comparison →](https://originchaindb.com/vs/pinecone)[OriginChainDB vs Qdrant Filtered vector search, payloads and where a second system of record appears. Read the comparison →](https://originchaindb.com/vs/qdrant)[OriginChainDB vs Weaviate Modules, hybrid search and the operational cost of a separate vector tier. Read the comparison →](https://originchaindb.com/vs/weaviate)[OriginChainDB vs Milvus Index variety and GPU acceleration versus one atomic multimodal commit. Read the comparison →](https://originchaindb.com/vs/milvus) operational stores [OriginChainDB vs MongoDB Documents with vector search versus rows, vectors, full-text and graph in one write. Read the comparison →](https://originchaindb.com/vs/mongodb)[OriginChainDB vs Aerospike Low-latency key-value at scale, compared on query shape and consistency. Read the comparison →](https://originchaindb.com/vs/aerospike) ## Not sure which page to read? Describe the workload. A working session with an engineer on your data shapes and query patterns. If another database is the better fit, we will say so. [Book a technical walkthrough](https://originchaindb.com/contact?topic=walkthrough) [Start free](https://app.originchaindb.com/signup) --- # Contact the OriginChainDB team | OriginChainDB Canonical source: https://originchaindb.com/contact Sitemap last modified: 2026-09-27T14:44:27.000Z Contact OriginChainDB # Start with what you want to build. Discuss a workload, work through a technical issue, or reach the security team directly. Three ways to reach us ## Choose a direct route. 1. ### Evaluate a workload Discuss your queries, deployment options and a technical walkthrough. [sales@originchain.ai](mailto:sales@originchain.ai) 2. ### Get technical help Include the error, request ID and a small reproducible example. Leave credentials out. [support@originchain.ai](mailto:support@originchain.ai) 3. ### Report a security issue Send a private vulnerability report to the security team. [security@originchain.ai](mailto:security@originchain.ai)[Read the disclosure guidance](https://originchaindb.com/security) ### For a useful workload discussion Include the data models you need, an approximate dataset size and the query you want to run. Add region or deployment preferences if you have them. [Explore the architecture first](https://originchaindb.com/architecture) Workload evaluation ## Tell us about it. Use the form to prepare an email to the sales team. Name Required Email Required Topic What are you working on? Required Opens your email app with a draft to sales@originchain.ai. You review and send it there. Prefer to write directly? [sales@originchain.ai](mailto:sales@originchain.ai) ## Keep exploring while you prepare. Start with the product, the implementation path, or the current service status. [Understand the database See how SQL, vector, graph and full-text fit together.](https://originchaindb.com/product/database)[Run a first query Follow the setup and query examples in the docs.](https://originchaindb.com/docs/quickstart)[Check service status Review the published status information.](https://originchaindb.com/status) --- # Example application workloads | OriginChainDB Canonical source: https://originchaindb.com/customers Sitemap last modified: 2026-09-23T08:33:23.000Z customers # Example application workloads Explore application patterns for OriginChainDB. These are workload examples; published customer references will be identified separately. 01 · what people run ## Three shapes we see every week. ### Retrieval for production assistants Documents, their embeddings and the full-text postings that catch exact terms live in one store. One write, one commit, one query path — no sync job between a vector index and the system of record. [RAG on one engine →](https://originchaindb.com/solutions/rag) ### Fraud and risk graphs Entities as rows, relationships as graph edges, behaviour as vectors — traversed and scored in the same request, inside the customer's region, with a per-request audit log. [Fraud detection →](https://originchaindb.com/solutions/fraud-detection) ### Personalisation at request time Catalogue rows, taste vectors and click history queried together under a latency budget, on a dedicated instance that scales compute and storage independently. [Personalisation →](https://originchaindb.com/solutions/personalization) ## Bring your workload. We'll measure it with you. A structured pilot on a dedicated instance, with success criteria agreed up front and a written report at the end. [Start a pilot](https://originchaindb.com/pilot) [Book a technical walkthrough](https://originchaindb.com/contact?topic=walkthrough) --- # Documentation | OriginChainDB Canonical source: https://originchaindb.com/docs Sitemap last modified: 2026-09-23T08:33:23.000Z ORIGINCHAINDB / DEVELOPERS # Documentation Set up your database, choose a query model, and follow a working example. [Run your first query](https://originchaindb.com/docs/quickstart)[Explore query guides](https://originchaindb.com/docs#docs-query-start)[API reference](https://originchaindb.com/docs/api) Queries / four data models ## Start with the question. Choose a path to its quickstart and implementation examples. One source record, different questionsInteractive illustration `knowledge.guides` ID guide:07 Title Battery replacement Product AX-7 Keep IDs with your source. SQL / structured records ### Which guides cover AX-7? Filter records by product, then select the fields you need. Example resultguide:07 · Battery replacement[Start with SQL](https://originchaindb.com/docs/sql/quickstart) Supply embeddings and configure the relevant indexes. Record, vector and full-text writes are separate; your application coordinates combined retrieval. Prefer natural language? [Try ASK](https://originchaindb.com/docs/nql/quickstart) as an optional query interface over your registered schemas. [## Use the console Explore records, run queries and inspect relationships.](https://originchaindb.com/docs/dashboard) [## Connect your application Choose an SDK, a database client or the HTTP API.](https://originchaindb.com/docs/sdk) ## Browse documentation Six sections. One place to start. [### Guides](https://originchaindb.com/docs) First steps and console walkthroughs. - [Quickstart](https://originchaindb.com/docs/quickstart) - [Console walkthrough](https://originchaindb.com/docs/dashboard) - [Example library](https://originchaindb.com/docs/examples) [### Data models](https://originchaindb.com/docs/schemas) Schemas, fields, records and relationships. - [Define a schema](https://originchaindb.com/docs/schemas/tutorial) - [Fields & types](https://originchaindb.com/docs/schemas/reference) - [Write records](https://originchaindb.com/docs/insert) [### Queries](https://originchaindb.com/docs/choosing) SQL, vector, graph and full-text queries. - [SQL](https://originchaindb.com/docs/sql) - [Vector](https://originchaindb.com/docs/vector) - [Graph](https://originchaindb.com/docs/graph) - [Full-text](https://originchaindb.com/docs/fts) [### Integrations](https://originchaindb.com/docs/sdk) SDKs, compatible clients and adapters. - [Client SDKs](https://originchaindb.com/docs/sdk) - [Database clients](https://originchaindb.com/docs/connect-wire-protocols) - [Elasticsearch adapter](https://originchaindb.com/docs/elasticsearch) [### Operations](https://originchaindb.com/docs/deploy) Deployments, security and troubleshooting. - [Deployment](https://originchaindb.com/docs/deploy) - [Authentication](https://originchaindb.com/docs/auth) - [Troubleshooting](https://originchaindb.com/docs/troubleshooting) [### API reference](https://originchaindb.com/docs/api) Endpoints, errors, limits and releases. - [HTTP endpoints](https://originchaindb.com/docs/api) - [Error codes](https://originchaindb.com/docs/errors) - [Limits](https://originchaindb.com/docs/rate-limits) Looking for an answer? Search the docs or ask a question about your workflow. [Ask the docs](https://originchaindb.com/docs/ai) --- # Ask AI - the OriginChainDB docs assistant Canonical source: https://originchaindb.com/docs/ai Sitemap last modified: 2026-09-14T18:28:08.000Z docs · ask AI BETA # Ask AI. Ask anything about OriginChainDB. Answers are grounded in the engine's actual grammar — schema fields, SQL surface, Cypher subset, vector + FTS runtime calls. Your conversation is stored in this browser tab only — survives navigation, lost on hard refresh. [Log in](https://app.originchain.ai/login?next=/docs/ai) to save chats across devices (coming soon). try one of these ⌘ / Ctrl + Enter to send --- # HTTP API - OriginChainDB AI-native database Canonical source: https://originchaindb.com/docs/api Sitemap last modified: 2026-09-13T20:18:50.000Z 03 · http api · 32 endpoints # HTTP API reference. One base URL per tenant instance. TLS 1.3 only. Bearer auth on every `/v1/...` path; mutating routes additionally honour `Idempotency-Key`. Request bodies cap at 8 MiB (the NDJSON batch route lifts that cap and applies its own per-line + total-buffer accounting). auth & headers | Header | Required | Notes | | --- | --- | --- | | Authorization: Bearer | Every /v1/* route | Tenant-scoped. /health, /ready, /metrics are public. | | Idempotency-Key: | Optional, mutating routes | Same key + same body = cached response. Different body with same key = 409. | | Content-Type | POST / PUT | `application/json` by default; `text/plain` for schemas; `application/x-ndjson` on the streaming batch route. | | Accept: text/event-stream | /v1/tenants/:t/watch | SSE stream. Server pushes one event per change burst. | | X-OC-Query-Id | Response only | ULID for the query. Pass it to `POST /v1/queries/:id/cancel`. | | X-OC-Replication: degraded | Response only | Write succeeded here, but the required follower acks did not arrive within the sync window. Nothing guarantees a replica holds the write, so a failover can lose it. Surface as a warning. | | Retry-After: | Response only (429) | Honour it. Clients should back off, not hammer. | errors Every non-2xx response is a JSON document of the form `{ "error": "code", "message": "...", "request_id": "..." }`. Quote `request_id` in support tickets. | Status | Code | Meaning | | --- | --- | --- | | 400 | validation_failed | Body or query parameters malformed. | | 401 | unauthorized | Bearer missing, invalid, or not scoped to this tenant. | | 402 | quota_exceeded | Authed and under RPS, but credit is exhausted. | | 403 | forbidden | Token cannot reach this resource. | | 404 | not_found | Schema / row / migration not registered. | | 409 | conflict | Idempotency replay-mismatch, lease busy, or migration wrong-state. | | 413 | body_too_large | Body over 8 MiB (or NDJSON line over 1 MiB). | | 429 | rate_limited | Per-bearer token bucket drained. Honour `Retry-After`. | | 499 | cancelled | Query cancelled via `POST /v1/queries/:id/cancel`. | | 500/503 | server_error / fenced | 5xx; retry with backoff if idempotent. 503 means this node is refusing writes - it no longer holds the writer lease, or its store is fenced and needs an operator. Re-read the lease to find the current holder. | ## Schemas TOML manifests describe a table - its primary key, columns, indexes, and relations. The engine indexes everything off the manifest, so registering one is the prerequisite to any row write. POST /v1/tenants/:tenant/schemas - Register or update a TOML manifest. Body is the raw TOML; Content-Type: text/plain. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/schemas" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: text/plain" \ --data-binary @orders.toml ``` #### TOML body ``` namespace = "trading" table = "orders" primary_key = ["order_id"] [[columns]] name = "order_id" ty = "str" required = true [[columns]] name = "symbol" ty = "str" [[columns]] name = "qty" ty = "i64" # Indexes are their own blocks - there is no "indexed = true" on columns. [[indexes]] name = "by_symbol" columns = ["symbol"] ``` response 200 OK ``` { "id": "trading.orders", "version": 1 } ``` GET /v1/tenants/:tenant/schemas - List every schema id registered for the tenant. request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/schemas" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` ["trading.orders", "trading.trades", "trading.users"] ``` GET /v1/tenants/:tenant/schemas/:id - Fetch the raw TOML for a schema id. Response is text/plain, not JSON. request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/schemas/trading.orders" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK · 404 if unknown ``` # text/plain namespace = "trading" table = "orders" primary_key = ["order_id"] ... ``` ## Rows (typed CRUD) Insert, batch, and read against a registered schema. Single-row writes are atomic; the batch endpoint accepts a JSON array (atomic in one commit) or NDJSON via `application/x-ndjson` (streamed in flushable chunks). POST /v1/tenants/:tenant/rows/:schema - Upsert a single row. `?expect=insert` skips the prior-state read for pure-insert bulk loads. Send `Idempotency-Key` to make retries safe. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/rows/trading.orders" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" \ -d '{ "order_id": "o-0001", "symbol": "AAPL", "qty": 100 }' ``` #### json body ``` { "order_id": "o-0001", "symbol": "AAPL", "qty": 100 } ``` response 200 OK · 400 validation · 409 idempotency replay-mismatch ``` { "ok": true, "lsn": { "segment": 4, "offset": 8421007 } } ``` POST /v1/tenants/:tenant/rows/:schema/_batch - Atomic batch (one durable commit). JSON body: a row array. NDJSON body (Content-Type: application/x-ndjson): streamed; flushes every `?chunk=N` rows (default 1000, max 10000). 8 MiB body cap is disabled on this route. request #### JSON array ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/rows/trading.orders/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: bulk-2026-05-01-batch-1" \ -d '[ { "order_id": "o-1", "symbol": "AAPL", "qty": 100 }, { "order_id": "o-2", "symbol": "MSFT", "qty": 250 } ]' ``` #### NDJSON stream ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/rows/trading.orders/_batch?chunk=2000" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/x-ndjson" \ --data-binary @orders.ndjson # orders.ndjson {"order_id":"o-1","symbol":"AAPL","qty":100} {"order_id":"o-2","symbol":"MSFT","qty":250} ``` response 200 OK ``` { "inserted": 2, "lsn": { "segment": 4, "offset": 8425112 } } ``` GET /v1/tenants/:tenant/rows/:schema/:pk - Read a row by its single-column primary key. Composite PKs must use POST /query with a ColumnScan plan. request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/rows/trading.orders/o-0001" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK · 404 not found ``` { "order_id": "o-0001", "symbol": "AAPL", "qty": 100, "_oc_row_version": 1 } ``` ## Query - Plan tree (JSON) The engine's native execution surface. POST a JSON Plan tree and get rows back. `?explain=true` returns the executed plan annotated with stats (EXPLAIN ANALYZE). Cancel an in-flight plan with the ULID handed back in `X-OC-Query-Id`. POST /v1/tenants/:tenant/query - Execute a Plan tree. Bare response is `Vec`; with `?explain=true` it's `{rows, explain}`. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/query" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "Limit": { "n": 10, "child": { "Filter": { "predicate": { "Eq": ["status", "pending"] }, "child": { "Scan": { "schema": "trading.orders" } } } } } }' ``` response 200 OK · 400 plan parse · 499 cancelled ``` [ { "order_id": "o-0001", "symbol": "AAPL", "status": "pending", ... }, ... ] ``` POST /v1/tenants/:tenant/query?explain=true - EXPLAIN ANALYZE: executes the plan and returns it annotated with per-node row counts and µs timings. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/query?explain=true" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "Scan": { "schema": "trading.orders" } }' ``` response 200 OK ``` { "rows": [ ... ], "explain": { "op": "Scan", "schema": "trading.orders", "rows_out": 412, "elapsed_us": 1340, "children": [] } } ``` POST /v1/queries/:id/cancel - Flip the cancellation token for an in-flight plan. The id is the ULID returned in `X-OC-Query-Id` on the original request. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/queries/01HW7G5...JZ/cancel" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK (always; cancelled=false if already finished) ``` { "cancelled": true } ``` GET /v1/tenants/:tenant/watch - Server-Sent Events stream. The connection holds open; the server pushes one event per change burst against the subscribed schemas. Ctrl-C / client close ends the subscription. request #### cURL ``` curl -N "$OC_BASE_URL/v1/tenants/$OC_TENANT/watch?schemas=trading.orders" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Accept: text/event-stream" ``` response 200 OK (text/event-stream) ``` event: snapshot data: { "rows": [ ... ] } event: snapshot data: { "rows": [ ... ] } ``` ## SQL A SQL surface over the same engine. SELECT, INSERT (+ RETURNING) and UPDATE execute; DELETE via /sql translates only (returns the pk, doesn't remove the row yet). BEGIN/COMMIT transactions and CREATE TABLE are supported too. Aggregates, OUTER JOINs, and chained 3+ table joins (up to 32) are supported. POST /v1/tenants/:tenant/sql - POST `{"sql": "..."}`. Response shape varies by `kind`: `select` -> `{kind, rows}`; `insert` -> `{kind, schema, rows}`; `delete` -> `{kind, schema, pk}`. request #### SELECT ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT order_id, symbol, qty FROM trading.orders WHERE status = '"'"'pending'"'"' LIMIT 10" }' ``` #### GROUP BY + HAVING ``` { "sql": "SELECT symbol, SUM(qty) AS shares FROM trading.orders GROUP BY symbol HAVING SUM(qty) > 1000" } ``` #### INNER JOIN (2) ``` { "sql": "SELECT u.email, o.order_id FROM trading.orders o INNER JOIN trading.users u ON o.user_id = u.user_id WHERE o.status = 'pending'" } ``` #### Chained JOIN (3+) ``` { "sql": "SELECT u.email, t.exchange, o.symbol FROM trading.orders o INNER JOIN trading.users u ON o.user_id = u.user_id INNER JOIN trading.trades t ON o.order_id = t.order_id" } ``` #### LEFT/RIGHT/FULL OUTER ``` { "sql": "SELECT u.email, o.order_id FROM trading.users u LEFT OUTER JOIN trading.orders o ON o.user_id = u.user_id" } ``` response 200 OK · 400 parse / unsupported ``` // SELECT { "kind": "select", "rows": [{"order_id":"o-1","symbol":"AAPL","qty":100}, ...] } // INSERT (re-issue against /rows/:schema with the returned rows) { "kind": "insert", "schema": "trading.orders", "rows": [...] } // DELETE (re-issue against /rows/:schema/:pk) { "kind": "delete", "schema": "trading.orders", "pk": "o-1" } ``` ## Vector search HNSW ANN with cosine / dot / L2 metrics and tunable speed/recall. Default high_recall mode hits recall@10 = 0.96 at 100k vectors with p99 109 ms; fast mode runs p99 37 ms at recall 0.69. The IVF-PQ index compresses each vector for large corpora; no measured result is published at 100M scale. Optional metadata is stored alongside each vector and queryable as an equality filter on topk. Brute-force fallback for small N. POST /v1/tenants/:tenant/vector/:table/put - Upsert one vector. Optional `metadata` object is indexed for filtered topk. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/vector/embeddings/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "doc-001", "embedding": [0.012, -0.443, ...], "dim": 384, "metric": "cosine", "metadata": { "lang": "en", "tier": "premium" } }' ``` response 201 Created ``` // 201 Created (no body) ``` POST /v1/tenants/:tenant/vector/:table/topk - k-nearest neighbour search. `mode=hnsw` (default) or `mode=bruteforce`. Non-empty `filter` triggers HNSW + post-filter on metadata equality. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/vector/embeddings/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.011, -0.439, ...], "k": 10, "dim": 384, "metric": "cosine", "mode": "hnsw", "filter": { "lang": "en" } }' ``` response 200 OK ``` [ { "id": "doc-001", "score": 0.973 }, { "id": "doc-127", "score": 0.952 }, ... ] ``` ## Full-text search Per-field, per-tenant inverted index. Tokenizer is UAX #29 (Latin / Cyrillic / CJK / Arabic / Hindi). Modes: `boolean` (AND), `bm25` (ranked, default Lucene k1=1.2 b=0.75), and `phrase` (exact contiguous tokens). POST /v1/tenants/:tenant/fts/:table/:field - Index a document under (table, field). Re-indexing the same `doc_id` cleans stale postings. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/fts/articles/body" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "art-001", "text": "OriginChain ships managed AI-native storage." }' ``` response 201 Created ``` // 201 Created (no body) ``` GET /v1/tenants/:tenant/fts/:table/:field?q=&mode=&k= - Search. `mode=boolean|bm25|phrase` (default boolean). `k` caps results in `mode=bm25` (default 10). request #### boolean ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/fts/articles/body?q=substrate%20managed&mode=boolean" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### bm25 ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/fts/articles/body?q=managed%20storage&mode=bm25&k=20" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### phrase ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/fts/articles/body?q=key-value%20storage&mode=phrase" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` // boolean / phrase ["art-001", "art-007"] // bm25 [ { "doc_id": "art-001", "score": 7.42 }, { "doc_id": "art-007", "score": 5.18 } ] ``` ## Elasticsearch API Point the official @elastic client - or any Elasticsearch 7.x REST client - at /v1/tenants/:tenant/es/. The cluster reports version 7.14.2. Documents commit with the row, so a write is searchable immediately, and row-level security and column masking apply to the query itself. Search also answers the filters aggregation and the term suggester, and a _bulk action may carry if_seq_no so a write lands only if nobody changed the document first. New here? Start with [Connect an Elasticsearch client](https://originchaindb.com/docs/connect-elasticsearch). PUT/v1/tenants/:tenant/es/:index- Create an index and its mapping. Field types are the Elasticsearch types your client already sends. request #### cURL ``` curl -X PUT "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @mapping.json ``` #### JSON body ``` { "mappings": { "properties": { "name": { "type": "text" }, "brand": { "type": "keyword" }, "price": { "type": "integer" } } } } ``` response 200 OK ``` { "acknowledged": true, "shards_acknowledged": true, "index": "shop.products" } ``` PUT/v1/tenants/:tenant/es/:index/_doc/:id- Index (create or replace) a document by _id. Searchable the instant the call returns - no refresh to wait on. request #### cURL ``` curl -X PUT "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_doc/sku-8842" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @doc.json ``` #### JSON body ``` { "name": "Carbon Marathon", "brand": "Aero", "price": 149 } ``` response 201 Created ``` { "_index": "shop.products", "_id": "sku-8842", "result": "created", "_version": 1 } ``` POST/v1/tenants/:tenant/es/_bulk- Bulk index / update / delete. NDJSON: an action line, then (for index/update) a source line. The fast path for ingest. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/_bulk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/x-ndjson" \ --data-binary @ops.ndjson ``` #### NDJSON body ``` {"index":{"_index":"shop.products","_id":"sku-1207"}} {"name":"Trail 24","brand":"Aero","price":89} {"update":{"_index":"shop.products","_id":"sku-8842"}} {"doc":{"price":139}} ``` response 200 OK ``` { "errors": false, "items": [ { "index": { "_id": "sku-1207", "status": 201 } }, { "update": { "_id": "sku-8842", "status": 200 } } ] } ``` POST/v1/tenants/:tenant/es/:index/_update/:id- Partial update. The fields in doc are merged in; every other field on the document is preserved. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_update/sku-8842" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @patch.json ``` #### JSON body ``` { "doc": { "price": 139 } } ``` response 200 OK ``` { "_index": "shop.products", "_id": "sku-8842", "result": "updated", "_version": 2 } ``` POST/v1/tenants/:tenant/es/:index/_search- Search with the Query DSL, with aggregations in the same request. RLS and column masking apply to the query itself. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_search" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @query.json ``` #### JSON body ``` { "query": { "bool": { "must": [ { "match": { "name": "marathon" } } ], "filter": [ { "term": { "brand": "Aero" } } ] } }, "aggs": { "by_brand": { "terms": { "field": "brand" } } } } ``` response 200 OK ``` { "took": 6, "hits": { "total": { "value": 1, "relation": "eq" }, "hits": [ { "_id": "sku-8842", "_score": 7.41, "_source": { "name": "Carbon Marathon", "brand": "Aero", "price": 149 } } ] }, "aggregations": { "by_brand": { "buckets": [ { "key": "Aero", "doc_count": 1 } ] } } } ``` POST/v1/tenants/:tenant/es/:index/_count- Count the documents matching a query. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_count" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @query.json ``` #### JSON body ``` { "query": { "match": { "name": "marathon" } } } ``` response 200 OK ``` { "count": 1 } ``` GET/v1/tenants/:tenant/es/:index/_doc/:id- Read one document by id. HEAD on the same path answers existence only, and /_source/:id returns the document with no envelope. request ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_doc/sku-8842" \ -H "Authorization: Bearer $OC_API_KEY" # bare document, one field curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_source/sku-8842?_source_includes=price" \ -H "Authorization: Bearer $OC_API_KEY" ``` response ``` { "_index": "shop.products", "_id": "sku-8842", "_version": 2, "found": true, "_source": { "name": "Carbon Marathon", "brand": "Aero", "price": 139 } } ``` GET/v1/tenants/:tenant/es/:index/_field_caps- What each field is, read from the stored mapping. This is the call dashboards make on connect. request ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_field_caps?fields=*" \ -H "Authorization: Bearer $OC_API_KEY" ``` response ``` { "indices": ["shop.products"], "fields": { "name": { "text": { "type": "text", "searchable": true, "aggregatable": false } }, "brand": { "keyword": { "type": "keyword", "searchable": true, "aggregatable": true } }, "price": { "long": { "type": "long", "searchable": true, "aggregatable": true } } } } ``` POST/v1/tenants/:tenant/es/:index/_delete_by_query- Delete matching documents. Honors max_docs; an unbounded whole-table delete is refused rather than run. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/es/shop.products/_delete_by_query" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @query.json ``` #### JSON body ``` { "query": { "term": { "brand": "Aero" } }, "max_docs": 100 } ``` response 200 OK ``` { "total": 3, "deleted": 3, "failures": [] } ``` ## Graph traversal Reads relations declared on a manifest. PKs are passed as path-encoded strings (single-column, string-typed). GET /v1/tenants/:tenant/graph/:schema/neighbors?rel=&pk= - Forward one-hop. Returns the destination PKs along `rel` from `pk`. request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/graph/social.users/neighbors?rel=follows&pk=alice" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` ["bob", "carol", "dave"] ``` GET /v1/tenants/:tenant/graph/:schema/reverse?rel=&pk= - Inbound one-hop: who points AT `pk` along `rel`. Works only when `from_table != to_table` (see oc-graph STATUS for the self-relation caveat). request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/graph/social.users/reverse?rel=follows&pk=bob" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` ["alice", "frank"] ``` GET /v1/tenants/:tenant/graph/:schema/bfs?rel=&pk=&max_depth= - Breadth-first search up to `max_depth` (default 3). Returns reachable nodes with depth. request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/graph/social.users/bfs?rel=follows&pk=alice&max_depth=2" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` [ { "pk": "bob", "depth": 1 }, { "pk": "carol", "depth": 1 }, { "pk": "ed", "depth": 2 } ] ``` GET /v1/tenants/:tenant/graph/:schema/path?rel=&src=&dst=&max_depth= - Reachability check: is there an `rel`-path from `src` to `dst` within `max_depth` hops? request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/graph/social.users/path?rel=follows&src=alice&dst=ed&max_depth=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` { "reachable": true } ``` GET /v1/tenants/:tenant/graph/:schema/dijkstra?rel=&src=&dst=&weights_json= - Weighted shortest-path. `weights_json` is a JSON object mapping `"|"` -> f64. A manifest weight-column variant is available - contact support. request #### cURL ``` curl --get "$OC_BASE_URL/v1/tenants/$OC_TENANT/graph/road.cities/dijkstra" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "rel=connects" \ --data-urlencode "src=NYC" \ --data-urlencode "dst=SFO" \ --data-urlencode 'weights_json={"NYC|CHI":2.0,"CHI|DEN":1.5,"DEN|SFO":1.2}' ``` response 200 OK ``` { "cost": 4.7 } ``` ## Online schema migrations Submit a diff, watch backfill progress, then cut over atomically. Aborts are allowed pre-cutover only. Every state transition is durably journaled. POST /v1/tenants/:tenant/migrations - Submit a migration. Body: `{schema, diff}` where `diff` is a `Vec`. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/migrations" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "schema": "trading.orders", "diff": [ { "AddColumn": { "name": "venue", "type": "str", "default": "NYSE" } } ] }' ``` response 200 OK ``` { "id": "0192ab...", "tenant": "01HW...ZZ", "schema": "trading.orders", "state": "Backfilling", "diff": [ ... ], "progress": { "rows_seen": 0, "rows_total": null } } ``` GET /v1/tenants/:tenant/migrations - List every migration the tenant has submitted (active + terminal). request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/migrations" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` [ { "id": "...", "schema": "trading.orders", "state": "Backfilling", ... }, ... ] ``` GET /v1/tenants/:tenant/migrations/:id - Read one migration. Poll for backfill progress. request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/migrations/0192ab..." \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK · 404 unknown ``` { "id": "...", "state": "ReadyToCutover", "progress": { "rows_seen": 1240000, ... } } ``` POST /v1/tenants/:tenant/migrations/:id/cutover - Atomic cutover. Only legal in `ReadyToCutover` state. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/migrations/0192ab.../cutover" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK · 409 wrong state ``` { "id": "...", "state": "Completed", ... } ``` POST /v1/tenants/:tenant/migrations/:id/abort - Abort. Only legal pre-cutover. Once the migration is `Completed`, abort returns 409. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/tenants/$OC_TENANT/migrations/0192ab.../abort" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK · 409 already cut over ``` { "id": "...", "state": "Aborted", ... } ``` GET /v1/tenants/:tenant/migrations/_audit - Append-only audit log of every state transition (submit / cutover / abort / auto-cutover) with actor + UNIX timestamp. request #### cURL ``` curl "$OC_BASE_URL/v1/tenants/$OC_TENANT/migrations/_audit" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` [ { "id": "...", "schema": "trading.orders", "state": "Backfilling", "at_secs": 1714512000, "actor": "...", "event": "submit" }, { "id": "...", "schema": "trading.orders", "state": "Completed", "at_secs": 1714512042, "actor": "...", "event": "cutover" } ] ``` ## Replication (admin) Lease-driven active-passive coordination. The lease holder is the sole writer; followers tail frames from the leader. These endpoints are operational, not application-facing - most tenants never call them. GET /v1/replication/lease - Read the current lease (or null if vacant). request #### cURL ``` curl "$OC_BASE_URL/v1/replication/lease" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` { "epoch": 4, "holder": "writer-a", "expires_at_secs": 1714512030, "etag": "..." } ``` POST /v1/replication/lease - Try-acquire the lease. `ttl_secs` defaults to 30. Returns 409 if held by another writer. request #### cURL ``` curl -X POST "$OC_BASE_URL/v1/replication/lease" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "holder": "writer-a", "ttl_secs": 30 }' ``` response 200 OK · 409 busy ``` { "epoch": 5, "holder": "writer-a", "expires_at_secs": 1714512060, "etag": "..." } ``` GET /v1/replication/frames?since_segment=&epoch= - Export the log from `since_segment` onwards as hex-encoded entries. Management-plane convenience; production followers stream raw bytes via the dedicated transport. request #### cURL ``` curl "$OC_BASE_URL/v1/replication/frames?since_segment=4&epoch=5" \ -H "Authorization: Bearer $OC_TOKEN" ``` response 200 OK ``` ["7b226c736e223a..."] ``` ## Health & observability All three are public (no auth) so load balancers and Prometheus scrapers can probe without a credential. GET /health - Liveness - process is up and the log is mounted. request #### cURL ``` curl "$OC_BASE_URL/health" ``` response 200 OK ``` { "status": "ok" } ``` GET /ready - Readiness - the bare call reports storage health and answers 200 even when `status` is degraded. Use `?require=write` (or `?require=read`) for the strict probe; `pressure` carries the reason. Follower lag is not part of the check. request #### cURL ``` curl "$OC_BASE_URL/ready" ``` response 200 OK · 503 not ready ``` { "write_ready": true, "status": "ready" } ``` GET /metrics - Prometheus exposition - request latencies, cache hits, log bytes, replication counts. request #### cURL ``` curl "$OC_BASE_URL/metrics" ``` response 200 OK (text/plain) ``` # HELP oc_query_latency_ms /v1/query end-to-end latency. # TYPE oc_query_latency_ms histogram oc_query_latency_ms_bucket{le="10"} 9218 oc_replication_frames_total 41702 oc_plan_cache_hits_total 41190 ... ``` SQL - scope today `NATURAL JOIN` and `?` bind-param syntax are not in scope today. `UNION`/`INTERSECT`/`EXCEPT`, CTEs (`WITH` / `WITH RECURSIVE`), `ALTER TABLE` and `$1` positional binding all ship. Correlated subqueries do ship: `WHERE EXISTS` / `NOT EXISTS`, `col IN (SELECT ...)` and correlated scalar comparisons all run. (`SELECT`/`INSERT`/`UPDATE`, `CREATE TABLE`, and `BEGIN`/`COMMIT` execute; `DELETE` via /sql translates only.) Use explicit `JOIN ... ON` for multi-table reads. --- # Natural language queries - OriginChainDB Ask Canonical source: https://originchaindb.com/docs/ask Sitemap last modified: 2026-09-23T08:33:23.000Z reference · ask # Natural language (Ask) Send a plain-English question to Ask and inspect the query plan and rows it returns. The compiler runs a deterministic rule grammar first. If the grammar can't resolve a phrase against the catalog, it falls back to whatever LLM is configured on the instance. The compiled plan is cached - repeat questions are answered without recompiling. ## 1. Ask a question. what this does Send a sentence. Get rows. The `schemas` field tells the compiler which tables to consider - omit it to allow every registered schema (slower because the catalog is larger). #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "last 10 orders where status is pending", "schemas": ["shop.orders"] }' ``` #### Python ``` result = db.ask( "last 10 orders where status is pending", schemas=["shop.orders"], ) for row in result["rows"]: print(row) ``` #### TypeScript ``` const result = await db.ask( "last 10 orders where status is pending", { schemas: ["shop.orders"] }, ); for (const row of result.rows) { console.log(row); } ``` #### Go ``` result, err := db.Ask(ctx, "last 10 orders where status is pending") if err != nil { /* handle */ } for _, row := range result.Rows { fmt.Println(row) } ``` request fields | Field | Required | Notes | | --- | --- | --- | | nl | yes | The question in plain English. Aim for one sentence - multi-step questions are best split. | | schemas | no | Array of schema names to consider. Omitting it allows all schemas but is slower for large catalogs. | | show_plan | no | When `true`, the response includes the executed plan tree. | common mistakes - Ambiguous column references. "Top customers by amount" doesn't say whether "amount" is per-order or aggregated. Be specific: "top customers by total order amount". - Not scoping schemas. If you know the question is only about orders, pass `schemas=["shop.orders"]` - the compiler is faster and less likely to pick the wrong table. - Sensitive data in questions. If you fall back to the LLM, the question gets sent to whatever LLM is configured. Don't put PII or secrets in the question. ## 2. Inspect the plan. what this does Set `show_plan: true` to see the query plan the compiler picked. Useful when you're going to call the same question in a hot path - you'll want to know whether it's hitting the right indexes. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "top 5 customers by total spend", "schemas": ["shop.orders", "shop.customers"], "show_plan": true }' ``` #### Python ``` # show_plan returns the executor plan alongside the rows. # Useful for verifying that the compiler hit the right indexes. result = db.ask( "top 5 customers by total spend", schemas=["shop.orders", "shop.customers"], ) print(result.get("plan")) print(result["rows"]) ``` response ``` { "rows": [ { "customer_id": "c_3", "total": 124900 }, { "customer_id": "c_7", "total": 98200 }, ... ], "cache": "miss", "plan": { "kind": "Aggregate", "by": ["customer_id"], ... } } ``` For the operator catalog (`Scan`, `Filter`, `HashJoin`, `Aggregate`, etc.), see the [query reference](https://originchaindb.com/docs/query). ## 3. The plan cache. Every response includes a `cache` field with one of three values: | Value | What happened | | --- | --- | | "hit" | Cached plan was reused. Fast. | | "miss" | First time seeing this question (or close to it). Compiled fresh, then cached. | | "skip" | Cache was bypassed, usually because schemas were modified since the entry was cached. | Schema migrations evict cached entries that referenced the affected tables. There's no manual eviction endpoint today - if you need to clear the cache, register the schema again (a no-op write counts as a migration). --- # Authentication - OriginChainDB bearer tokens Canonical source: https://originchaindb.com/docs/auth Sitemap last modified: 2026-09-13T20:18:50.000Z reference · auth # Authentication Every OriginChainDB API call is authenticated with a bearer token in the `Authorization` header. There are no API keys, no signed-request schemes, no OAuth dance. One header, one token, every call. ## 1. Get a token. Tokens are created in the dashboard. Open [app.originchain.ai](https://app.originchain.ai/) → your instance → API tokens → Create token. - The token is shown once at creation - never again. Paste it into your secrets manager before closing the page. - The token starts with `oc_live_` (production) or `oc_test_` (free databases). - Tokens are scoped to a single instance. Multiple instances need multiple tokens. - You can create as many tokens as you like. Name them by use case ("github-actions", "etl-pipeline", "laptop-debug") so revoking the right one is easy. ## 2. Use the token. Set it in your shell, then attach it to every API call. The SDKs read it once at client construction and reuse it for the lifetime of the client. #### cURL ``` # Every API call includes one header: curl "https://$OC_HOST/v1/tenants/$OC_TENANT/schemas" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` from originchain import OriginChain db = OriginChain( base_url=os.environ["OC_BASE_URL"], bearer=os.environ["OC_BEARER"], tenant=os.environ["OC_TENANT"], ) ``` #### TypeScript ``` import { OriginChainClient } from "@originchain/sdk"; const db = new OriginChainClient({ baseUrl: process.env.OC_BASE_URL!, bearer: process.env.OC_BEARER!, }); ``` #### Go ``` db := originchain.NewClient(originchain.Config{ BaseURL: os.Getenv("OC_BASE_URL"), Bearer: os.Getenv("OC_BEARER"), }) ``` what happens when auth fails | Status | Cause | | --- | --- | | 401 unauthorized | No `Authorization` header, malformed header, or the token doesn't match any known token for the instance. | | 403 forbidden | Token is valid but targets a different instance. Check the endpoint hostname matches the token's instance. | | 404 instance not found | The endpoint URL is wrong, or the instance was deleted. | ## 3. Storing tokens safely. A bearer token grants full access to the instance it was created for. Treat it like a password. do - Store tokens in a secrets manager (1Password, Doppler, Infisical, or your platform's secrets manager). - Load them into apps via environment variables. - Create one token per use case so revoking is targeted. - Use short-lived tokens for CI/CD - rotate them on a schedule. don't - Commit tokens to git, even in private repos. - Send tokens to client-side code (browsers, mobile apps). They're full-access; anyone who inspects the JS sees the token. - Share tokens via Slack, email, or screenshots. - Put tokens in URLs - they end up in proxy logs, browser history, and shell history. ## 4. Rotation. Rotate tokens regularly even if nothing has gone wrong. The pattern is overlap, then revoke: 1. Create a new token in the dashboard. 2. Update your secret manager / env var with the new token. 3. Restart your app, or wait for the rolling deploy to pick it up. 4. Confirm requests are still landing (check the dashboard's recent activity panel). 5. Revoke the old token in the dashboard. Revocation is immediate. After you revoke a token, every request using it returns `401` within a few seconds. ## 5. If a token leaks. If you suspect a token was exposed (committed to a public repo, posted in a chat, found in a log file someone else can read): 1. Revoke the token immediately in the dashboard. This is the only action that matters. 2. Check the instance's activity panel for any requests you didn't make. 3. Create a new token and update your apps. 4. If the leak was a public repo, also [scrub the git history](https://docs.github.com/en/code-security/secret-scanning/about-secret-scanning) - the old commits still contain the token even if you delete the file. If you discover suspicious activity, email [security@originchain.ai](mailto:security@originchain.ai). We'll help you scope the exposure and check our logs. --- # Choosing a query shape - OriginChainDB Canonical source: https://originchaindb.com/docs/choosing Sitemap last modified: 2026-09-23T08:33:23.000Z decision guide # Choosing a query shape Choose the query surface that matches your task: structured filters, similarity, keywords, relationships, or a natural-language question. If your query looks like X, reach for Y. The capability matrix at the bottom shows what the engine does today, not what is planned. ## One store, every query shape. Declare the shapes a table needs — SQL columns, a vector field, a full-text field, graph edges — and one write keeps them all in sync. There is no ETL, no second database, and no index-refresh lag: the row and every index it feeds land in the same atomic commit. Each query shape is just a different lens over that one copy of the data. ONE WRITE · one atomic commit ↓ One store · your table rows · k/vvector indexinverted index · postingsgraph edges every shape reads the same data ↓ SQL rows + indexes → joins, filters, aggregates Vector kNN over the ANN index (HNSW / IVF-PQ) Full-text BM25 over the postings Graph traverse the edges (BFS, paths, PageRank) Natural language an LLM plans one of the shapes above ### How each shape executes. The exact path a query takes, from what you send to the rows that come back. SQL SQL text→parse & plan→index seek / scan→filter · join · group→result rows Vector query embedding→ANN walk (HNSW / IVF-PQ)→top-k by cosine / dot / L2→fetch rows Full-text query text→analyze (tokenize · stem)→postings · BM25→ranked doc_ids→fetch rows Graph start node(s)→walk edges (adjacency)→BFS · shortest path · PageRank→node set Natural language prompt→LLM plans a shape (SQL / Cypher)→execute→answer + source rows The write path write a row→one WAL commit→rows + vector + postings + edges→queryable at once ## The shapes, side by side. SQL `POST /v1/tenants/:t/sql` use for Filtering by exact-match WHERE on indexed columns. JOINs across up to 32 tables. GROUP BY with COUNT / SUM / AVG / MIN / MAX. LIMIT-bounded reads. avoid for Semantic similarity (use vector). Approximate string matching (use full-text). Multi-hop relationship questions (use graph). Natural-language questions from non-engineers (use Ask). recognise it by A Postgres-shaped query you'd write in DBeaver. Vector search `POST /v1/tenants/:t/vector/:table/topk` use for Semantic similarity. Cross-language retrieval. 'Find rows like this one' even when the keywords don't match. Recommendations. De-duplication by meaning. RAG retrieval before an LLM call. avoid for Exact identifiers, SKUs, error codes (use SQL or full-text). Structural traversal (use graph). recognise it by An embedding vector plus k. Full-text (BM25) `GET /v1/tenants/:t/fts/:table/:field` use for Exact phrase matching. Acronyms and product codes. Long-tail queries with unusual terms. Recall on documents containing the literal keyword. avoid for Conceptual queries where the user's wording doesn't match the document's wording (vector wins). Structured-field filtering (SQL is cheaper). recognise it by Words a human types into a search box. Graph traversal `GET/POST /v1/tenants/:t/graph/:schema/:algo` use for Multi-hop relationship questions ('orders from customers I haven't reviewed yet'). Social-graph walks. Dependency chains. Shortest path. PageRank or centrality analytics. Reachability checks. avoid for Single-table lookups (use SQL). Semantic similarity (use vector). Large-result analytics that aren't relationship-shaped (SQL is cheaper). recognise it by Multi-hop or path query: 'shortest path through citations'. Hybrid (vector + BM25) `Run vector topk + FTS in parallel, fuse client-side` use for Production retrieval. Catches both semantic match and literal keyword match. Generally outperforms either alone on standard benchmarks. avoid for Anything where one mode is structurally enough - don't fuse if vector alone is already perfect. recognise it by RAG retrieval before the LLM call. Natural language (Ask) `POST /v1/tenants/:t/ask` use for Non-technical users asking questions of structured data. Internal dashboards. Customer-support agents. Prototype-grade analytics without writing SQL. avoid for Latency-critical hot paths (a cold compile costs an LLM round-trip). Queries that need to be auditable to a single SQL string. recognise it by An English sentence. ## If your query looks like this... Pattern-match the left column against what you're trying to do; the middle and right columns are the answer. your queryreach forwhy `WHERE id = 'sku-9281'` `SQL` Exact lookup on an indexed primary key. `WHERE status = 'pending' AND amount_cents > 100` `SQL` Multi-predicate WHERE with AND on indexed columns. `Per-customer totals over paid orders` `SQL` GROUP BY customer + SUM(amount_cents). `Products similar to this one (no shared keywords)` `Vector` Semantic similarity via embedding distance. `Products described as 'lightweight running shoes for marathons'` `Hybrid (vector + BM25)` Catches the semantic match AND the literal keyword. `Find SKU ABC-1234-XL` `Full-text or SQL` Exact-token retrieval. Vector would dilute it semantically. `Path between paper A and paper Z through citations` `Graph (BFS / path)` Multi-hop walk. Cap with max_depth. `Shortest commute between two stations` `Graph (Dijkstra)` Weighted shortest path. Supply edge weights via the JSON weights map. `Most influential nodes in a network` `Graph (PageRank)` Iterative influence over a seed node set. `Customers in segment X (English question)` `Ask` Translates the sentence to a Plan against your schemas. ## Capability matrix. What's supported today. `yes` = works · `partial` = limited shape (see the relevant reference page) · `—` = not the right tool for this shape, or not yet supported. "Ask" inherits SQL's surface where the compiler can build the right Plan - if SQL doesn't support a construct (like HAVING or window functions), Ask can't either. | Feature | SQL | Vector | FTS | Graph | Ask | | --- | --- | --- | --- | --- | --- | | Exact-match WHERE on indexed col | yes | — | yes | — | yes | | AND-combined WHERE conditions | yes | — | — | — | yes | | OR in WHERE | — | — | — | — | partial | | IN (literal list) | yes | — | — | — | yes | | BETWEEN, IS NULL, LIKE | yes | — | — | — | yes | | GROUP BY + COUNT / SUM / AVG / MIN / MAX | yes | — | — | — | yes | | HAVING | — | — | — | — | partial | | ORDER BY | — | yes | yes | — | partial | | INNER / LEFT / RIGHT / FULL OUTER JOIN | yes | — | — | — | yes | | LIMIT | yes | yes | yes | yes | yes | | Uncorrelated IN (SELECT ...) | partial | — | — | — | partial | | Correlated subqueries / EXISTS | — | — | — | — | — | | CTEs (WITH) | — | — | — | — | — | | Window functions | — | — | — | — | — | | EXPLAIN | yes | — | — | — | yes | | Transactions (BEGIN/COMMIT/ROLLBACK) | yes | — | — | — | — | | HNSW · cosine / dot / L2 / Manhattan | — | yes | — | — | — | | IVF / IVF-PQ for 10M+ corpora | — | yes | — | — | — | | Metadata equality filter on topk | — | yes | — | — | — | | fast / high_recall mode selector | — | yes | — | — | — | | BM25 ranking | — | — | yes | — | — | | Boolean AND | — | — | yes | — | — | | Phrase (exact word order) | — | — | yes | — | — | | Fuzzy / typo tolerance | — | — | yes | — | — | | 18-language stemming · 9-language lemmas | — | — | yes | — | — | | Neighbours (forward + reverse) | — | — | — | yes | — | | BFS, path, all simple paths | — | — | — | yes | — | | Dijkstra / k-shortest weighted paths | — | — | — | yes | — | | PageRank, betweenness, eigenvector | — | — | — | yes | — | | Louvain, label-prop, components | — | — | — | yes | — | | Node2Vec / GraphSAGE embeddings | — | — | — | yes | — | | Atomic write (row + indexes + edges) | yes | yes | yes | yes | — | | Idempotency keys | yes | yes | yes | yes | yes | ## Combining shapes in one app. The shapes aren't mutually exclusive. A single bearer token can hit any of them, and the engine keeps row writes + vector + full-text indexes consistent. Three common patterns: - RAG retrieval before an LLM call. Run vector + full-text in parallel, fuse with Reciprocal Rank Fusion in your app, pass the top-k rows into your prompt. The rows you retrieve are from the same store as your application data, so authorization is consistent. - Graph filter + SQL projection. Use a multi-hop traversal to find candidate primary keys, then do a SQL `WHERE id IN (...)` for the full projection. Graph hop costs ~tens of ms; SQL projection is sub-millisecond. - Ask with show_plan. Let a non-technical user write the question, return the compiled plan, paste the equivalent SQL into your codebase. Ask becomes the prototype; the SQL becomes the production query. --- # MySQL wire protocol compatibility - OriginChainDB Canonical source: https://originchaindb.com/docs/compatibility/mysql-wire Sitemap last modified: 2026-09-13T20:18:50.000Z reference · mysql wire # Compatible with the MySQL wire protocol An OriginChainDB instance can expose a listener that is compatible with the MySQL wire protocol, so a stock MySQL client or driver connects, authenticates and runs SQL without a shim. It reads the same store your [HTTP API](https://originchaindb.com/docs/api) calls and your [PostgreSQL-wire](https://originchaindb.com/docs/connect-sql-client) sessions see - there is one copy of your data, not a replica. Whether it also writes depends on which [credential model](https://originchaindb.com/docs/compatibility/mysql-wire#credentials) your listener runs: a shared-credential listener reads and writes; a per-user listener is read-only, and refuses every write. off by default · enabled per instance The adapter is built into the engine but bound to nothing. No instance gets a MySQL-wire listener by default and enabling one is not self-serve today: it needs an address, a tenant, a credential and TLS material, so [talk to us](https://originchaindb.com/contact) and we will tell you whether your instance can be served at all - the [credential models](https://originchaindb.com/docs/compatibility/mysql-wire#credentials) section explains the one case where it cannot. Everything on this page also works today over the [SQL endpoint of the HTTP API](https://originchaindb.com/docs/sql), which needs no setup. ## What this is, and what it is not This is a compatibility statement about a protocol, not a database product. OriginChainDB implements the MySQL wire protocol independently so that clients written for it can talk to an OriginChainDB instance. We do not run, embed, fork or redistribute MySQL server software, and an instance is not a drop-in replacement for one: the SQL surface is OriginChainDB's, and the differences are listed under [limits](https://originchaindb.com/docs/compatibility/mysql-wire#limits). Concretely, the listener speaks protocol version 10: the standard handshake, the text protocol for ordinary queries, and binary prepared statements (`PREPARE` / `EXECUTE` / `CLOSE` / `RESET`, with `?` placeholders bound server-side). Statements arrive in MySQL's dialect and are translated into the engine's, construct by construct - never guessed at. Anything the translator does not recognise comes back as a clean error naming the construct, so a query either runs or fails; it is never quietly turned into a different query. Row-level security and column masking are applied inside the engine, below the wire, so a MySQL session sees exactly the rows and columns your policies allow - the same ones an HTTP call under that identity would see. ## Connect First find out whether your instance has a listener at all. Ask the capabilities endpoint: ``` curl -s -H "Authorization: Bearer $OC_TOKEN" \ "https://$OC_HOST/v1/capabilities" | jq '.wire.mysql' # Until a listener is configured for your instance: # { "available": false } # # "available" reports that CONFIGURATION IS COMPLETE. It is a precondition # for starting a listener, not an observation of a bound socket - so when it # turns true, still check the port and a real login before you rely on it. ``` Once it is enabled, connect the way you would connect to anything speaking this protocol: ``` # The port follows the convention MySQL clients expect. The host and port # your instance is actually published on are shown in the console once the # listener is enabled - use those, not this default. mysql --host "$OC_MYSQL_HOST" --port 3306 \ --user "$OC_DB_USER" --password \ --ssl-mode=REQUIRED \ --database "$OC_DB" ``` The authentication plugin depends on which credential model your listener runs: `mysql_native_password` for a shared credential, `caching_sha2_password` for per-user login. Both are what a stock client already negotiates, so you do not configure the plugin - you configure the model, and the plugin follows. ## TLS, and why a certificate is not optional The listener implements the standard TLS upgrade: it advertises TLS support in the handshake, a capable client answers by asking to upgrade before it authenticates, and the rest of the session runs inside TLS. That is exactly what `--ssl-mode=REQUIRED` does, so stock clients interoperate without special handling. a shared-credential listener fails open Configured with a shared credential and no certificate, the listener still binds and still serves - with TLS simply off. It does not refuse to start, and a client that does not ask to upgrade is answered in cleartext: passwords, queries, and every result row, including rows your policies filtered and columns they masked, cross the network in the clear. Nothing warns you at connect time, because from the client's point of view it worked. Never enable a shared-credential listener without a certificate, and set the option that refuses plaintext authentication rather than relying on every client to opt in. Per-user login is the opposite: it cannot be configured without TLS. That model's authentication delivers the password to the server in cleartext inside the session, so the engine forces the refuse-plaintext option on and will not start a per-user listener that has no certificate. If the tenant you are enabling must use per-user login - see below - TLS is settled for you. ## Credential models, and which one you will be on There are two, and the choice is not a preference - it is decided by whether your tenant has any database authorization to enforce: a database RBAC grant or a row-level security policy. Either one is enough; a row-security policy is not the lighter case. | Model | What the client sends | You are on it when | | --- | --- | --- | | shared | One listener-wide password. Every session arrives as the same identity. | Your tenant has no database RBAC grants and no row-level security policies at all. A shared credential cannot express a per-user grant or evaluate a per-user row policy, so the engine refuses to combine the two rather than serve them loosely. | | per-user | Each database user authenticates with its own password, and its own grants apply for the session. | Your tenant has any database RBAC grant, or any row-level security policy. This is the only model that can serve it. TLS is mandatory here, and the session is read-only. | a per-user listener is read-only A per-user MySQL session refuses every write. Not some writes, and not writes your grants happen to disallow - the listener marks the session read-only when it admits an authenticated identity, before any grant is consulted, because there is no write-parity bridge on this adapter. `INSERT`, `UPDATE`, `DELETE` and DDL all come back refused, and no configuration turns this on. Read that together with the table above and the consequence is the one to plan around: a tenant with a single RBAC grant or a single row-level security policy cannot write over the MySQL wire at all. That tenant must be on per-user login, and per-user login does not write. There is no combination of settings that yields an authorizing, writing MySQL session today. So the write examples on this page and under [MySQL examples](https://originchaindb.com/docs/examples/mysql) - CRUD, transactions and `ON DUPLICATE KEY UPDATE` - apply to a shared-credential listener only. If your tenant has any authorization to enforce, write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client), and keep the MySQL listener for reads. Put the two rules together and there is a real state that neither model can serve: a tenant that has database RBAC or a row-security policy but has no database user carrying a password. Shared login is refused because that authorization is present; per-user login is refused because there is no account to admit. That is the ordinary state of a tenant whose access has only ever been through API tokens, so it is worth checking before you plan the work. Creating one enabled database user with a password is the whole fix. ## A worked example: Sequelize over mysql2 This is the client we drove against the shipping engine - Sequelize 6.37 on top of mysql2 3.24 - so what follows is a transcript of what worked, not a sketch of what should. The shape to take away: declare your tables in code and the ORM drives a complete single-table workload; let it introspect the database instead and it will not get far. ### Connect and define ``` import mysql2 from "mysql2"; import { Sequelize, DataTypes } from "sequelize"; // Verified against the shipping engine with Sequelize 6.37 over mysql2 3.24. const sequelize = new Sequelize(process.env.OC_DB, process.env.OC_DB_USER, process.env.OC_DB_PASSWORD, { host: process.env.OC_MYSQL_HOST, port: 3306, dialect: "mysql", dialectModule: mysql2, // Ask for TLS explicitly. A listener with no certificate will happily // serve you in cleartext instead of refusing - see "TLS" above. dialectOptions: { ssl: { minVersion: "TLSv1.2", rejectUnauthorized: true } }, logging: false, }); // Describe the table in code. Do NOT call sequelize.sync() and do not rely on // queryInterface.describeTable() - both introspect with SHOW, which is refused. const Order = sequelize.define( "Order", { id: { type: DataTypes.STRING, primaryKey: true }, customer: DataTypes.STRING, total_cents: DataTypes.INTEGER, status: DataTypes.STRING, }, { tableName: "orders", timestamps: false, freezeTableName: true }, ); await sequelize.authenticate(); ``` ### Insert, select, update, delete ``` // INSERT await Order.create({ id: "o_1", customer: "c_7", total_cents: 4200, status: "open" }); // SELECT with a filter and an ordering const open = await Order.findAll({ where: { status: "open" }, order: [["total_cents", "DESC"]], limit: 10, }); // UPDATE await Order.update({ status: "shipped" }, { where: { id: "o_1" } }); // DELETE, bounded by a WHERE await Order.destroy({ where: { id: "o_1" } }); ``` ### Transactions Transactions are real, not acknowledged-and-ignored. A block that commits persists, and a block that rolls back leaves nothing behind - both were exercised. ``` // Both outcomes were exercised against the engine: a transaction that // commits, and one that rolls back leaving no trace. const tx = await sequelize.transaction(); try { await Order.create({ id: "o_2", customer: "c_7", total_cents: 900, status: "open" }, { transaction: tx }); await Order.update({ status: "paid" }, { where: { id: "o_2" }, transaction: tx }); await tx.commit(); } catch (err) { await tx.rollback(); throw err; } ``` More recipes, one per page, are under [MySQL examples](https://originchaindb.com/docs/examples/mysql). ## The catalog surface This is the section that decides whether a tool works, so read it before you adopt one. Metadata is split down the middle: the `information_schema` views are answered, and every `SHOW` form is refused. ### Answered: the information_schema views Five views are served, by the same shared catalog the other wire protocols read: `tables`, `columns`, `table_constraints`, `key_column_usage` and `referential_constraints`. They are ordinary queries, so they work from any client: ``` -- Answered by the shared catalog, the same one the other wire -- protocols and the HTTP API read. SELECT table_name FROM information_schema.tables WHERE table_schema = 'shop'; SELECT column_name, data_type FROM information_schema.columns WHERE table_name = 'orders'; ``` ### Refused: every SHOW form ``` mysql> SHOW TABLES; ERROR 1235 (42000): MySQL `SHOW …` metadata commands are not supported — the engine keeps no MySQL information schema. -- Refusing is the deliberate choice. The shared front door carries a -- PostgreSQL "SHOW " shim that would otherwise answer ANY SHOW with a -- single fabricated row reading x = 'on' - so SHOW TABLES would return a -- confident, wrong answer instead of an error. ``` `SHOW TABLES`, `SHOW DATABASES`, `SHOW COLUMNS`, `SHOW CREATE TABLE`, `SHOW INDEX`, `SHOW WARNINGS` and `SHOW VARIABLES LIKE …` all return an error. We would rather hand you an error you can act on than a fabricated row you would believe. Session variables are handled separately and honestly: the ones a connector reads at startup are answered with their true values, and one that does not exist is refused with the unknown-variable error rather than a made-up value. `SET NAMES` is accepted as a no-op, because every official connector opens with one. ## Limits Every one of these returns a clear error naming the construct. None of them silently returns a different answer. the one that decides adoption Schema reflection is split, and associations do not load. An ORM pointed at this listener can connect, authenticate and run a full single-table workload including transactions - we measured exactly that - but it cannot reflect a schema it was not told about, and it cannot follow relations between tables. So it suits an application whose tables you declare in code, and it does not suit one that models a domain with associations. If your application depends on relations, use the [SQL endpoint](https://originchaindb.com/docs/sql) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client) instead. Improving this is work in progress, and work in progress is not a feature - plan against what is written here. | Not supported | Why, and what to write instead | | --- | --- | | SHOW … | No MySQL information schema is kept. Query the `information_schema` views listed above instead. | | REPLACE INTO | Its delete-then-insert semantics reset unmentioned columns to defaults, which is not what the engine's upsert does. Use `INSERT … ON DUPLICATE KEY UPDATE`, which is translated. | | INSERT IGNORE | Silently swallowing errors is not a behaviour we will emulate. Handle the conflict explicitly. | | Multi-table DML | `DELETE t1, t2 FROM …`, `DELETE … USING …`, `DELETE … ORDER BY / LIMIT` and joined `UPDATE`s. The engine models single-table DML; a `DELETE` bounded by `WHERE` works. | | GROUP_CONCAT, SUBSTRING_INDEX | No engine equivalent to map them onto. Most other common scalar functions are translated or pass through unchanged. | | User variables (@x) | The engine has nowhere to keep session-scoped user variables. Keep the value in your application. | | Multi-statement queries | Send one statement per query. Do not enable your driver's multi-statement option. | | Double-quoted strings | `"x"` is parsed as an identifier, not a string. Spell strings with single quotes - `'x'` - which is what clients do anyway. | | Chunked parameter sends | A prepared statement whose parameter was streamed in chunks is refused rather than executed with an incomplete value. Bind the value in one go. | ## Trademark MySQL is a registered trademark of Oracle Corporation and/or its affiliates. OriginChainDB is not affiliated with, endorsed by, or sponsored by Oracle Corporation. OriginChainDB implements a compatible wire protocol so that existing MySQL clients can connect to it; it does not distribute MySQL software. --- # Oracle SQL dialect - what ports, what doesn't Canonical source: https://originchaindb.com/docs/compatibility/oracle-dialect Sitemap last modified: 2026-09-13T20:18:50.000Z reference · oracle sql dialect # Oracle SQL dialect OriginChainDB answers a documented subset of the Oracle SQL dialect — sequences, NVL, FROM DUAL, MINUS, CONNECT BY and more — so that SQL written for Oracle can run here largely unchanged. It is reachable through an ordinary PostgreSQL driver, over the same doors every other SQL client uses. read this first · the dialect, not the transport We implement the Oracle SQL dialect. We do not implement Oracle’s network protocol, and it is not planned — a decision, not a backlog item. There is no TNS listener, no Oracle Net negotiation and no OCI surface, so an Oracle client library cannot connect at all. SQL*Plus, SQL Developer, Oracle Instant Client, OCI, ODP.NET, the Oracle JDBC driver and python-oracledb in thick mode all speak Oracle Net. None of them will reach an instance, and no connection setting changes that. What ports is your SQL: point a PostgreSQL driver at the database and keep the statements. The dialect travels; the transport does not. Your statements arrive over a [PostgreSQL driver](https://originchaindb.com/docs/connect-sql-client) or the [HTTP /sql endpoint](https://originchaindb.com/docs/api#sql) — never over Oracle Net. ## How to connect. There is no Oracle-specific endpoint and nothing to switch on for the dialect. You connect exactly as any other SQL client does, and the Oracle spellings are understood on arrival. Two doors, same behaviour: - The PostgreSQL wire protocol — port 5432, TLS required, SCRAM-SHA-256. Any PostgreSQL driver works: JDBC’s PostgreSQL driver, psycopg, node-postgres, pgx, libpq, the PostgreSQL ODBC driver. Setup and per-client walkthroughs are on [Connect a SQL client](https://originchaindb.com/docs/connect-sql-client). - The HTTP /sql endpoint — one bearer token, no driver at all. See the [SQL reference](https://originchaindb.com/docs/sql). The dialect work happens in the SQL translator, above the door, on the one seam both surfaces share — so a statement that is accepted over HTTP is accepted over the wire, and refused the same way. PostgreSQL wire access is in limited preview This paragraph is about the PostgreSQL wire protocol — the only wire this page’s Oracle dialect is served over. PostgreSQL-wire access is enabled gradually, and self-serve enablement covers single-node instances today. If your instance has no SQL access panel in the console yet, the PostgreSQL listener is not switched on for you — everything on this page still works over the HTTP /sql endpoint. Do not read that across to the other wire protocols. Enablement of the MySQL- and SQL Server-compatible listeners is not self-serve today and is arranged with us; neither serves the Oracle dialect described here. ``` # A PostgreSQL client, talking Oracle SQL. Nothing Oracle-specific here. psql "host= port=5432 dbname=postgres \ user= password= sslmode=require" => SELECT NVL(region, 'unknown') AS region FROM sales.orders MINUS SELECT NVL(region, 'unknown') FROM sales.returns; ``` ## What the three tiers mean. The failure that matters in a dialect is not a refusal — it is a construct that is accepted and answered differently than Oracle would answer it, because you then get a wrong number with a 200 next to it. The engine is built to refuse loudly instead. So: - Generally available — shipped, and Oracle-equivalent within the scope named in the row. - Preview — shipped and reachable, but deliberately narrower than Oracle. The served shape is described; everything outside it is refused with a 400 that names the construct, never answered approximately. - Not supported — not accepted today. The rewrite that gets you the same answer is in the row. A boundary in the “preview” column is a real boundary, not a caveat for form’s sake. Read the row before you port a statement that depends on it. ## Generally available. These behave as Oracle does within the scope described. Two dialect-wide differences apply throughout and are worth internalising once: a zero-length string is not NULL here (Oracle treats '' as NULL), and a chosen operand is returned with its own type rather than coerced to the first argument’s type — so keep the operands of one call in one type. | Construct | What ships, and its scope | | --- | --- | | CREATE SEQUENCE nextval() / currval() | `CREATE SEQUENCE [IF NOT EXISTS] name [START WITH n] [INCREMENT BY n]`, then `nextval('name')` / `currval('name')` in an INSERT/UPDATE value position or as a bare FROM-less SELECT. The counter is durable before use — flushed to disk before the value is handed out — so a crash never re-issues an id, and a refused statement burns no value. Refused at CREATE, each by name: `CYCLE`, `MINVALUE`/`MAXVALUE`, `OWNED BY`, `AS `, `INCREMENT BY 0`. Oracle’s `CACHE`/`NOCACHE` has no analogue: every value here is durable, which is the `NOCACHE` behaviour. | | seq.NEXTVAL seq.CURRVAL | The Oracle pseudocolumn spelling, any case, routing to the same durable counter as the function spelling. It inherits every restriction of that spelling deliberately — same permitted call sites, same read-only and replica gating. `INSERT ... SELECT seq.NEXTVAL FROM t` is therefore still refused: that is the table-scanning case, and per-row sequence generation is not what this surface does. | | NVL(a, b) | Yields `a` when it is not NULL, else `b` — the same result as `COALESCE(a, b)`, which also ships. Wrong arity is refused at translate time rather than answered NULL, because a NULL is indistinguishable from a legitimate result. Carries the two dialect-wide differences above. | | NVL2(a, b, c) | Yields `b` when `a` is not NULL, else `c`; only the branch actually taken is evaluated. `CASE WHEN a IS NOT NULL THEN b ELSE c END` is the portable spelling and is also served. | | MINUS | Oracle’s `MINUS` and standard `EXCEPT` are the same operator, so the word is rewritten and the statement is then the `EXCEPT` the set-operation path has always served — it adds no narrowing of its own. The rewrite applies only outside string literals, quoted identifiers and comments, so `SELECT 'a MINUS b'` keeps its value and a column named `"MINUS"` keeps its name. | | FROM DUAL | The clause is recognised on a single un-joined, un-aliased factor named `dual` or `sys.dual` (any case) and stripped, so `SELECT 1 FROM dual` behaves exactly like the FROM-less spelling — constants, arithmetic, scalar functions, CASE, CAST, and WHERE / ORDER BY / LIMIT over the one synthetic row. The residual: recognition is textual, so a table genuinely named `dual` is reached only by shapes that read table content (a column reference, `*`, GROUP BY, a join, an alias). Avoid naming a table `dual`. | ## Preview. Shipped and reachable, narrower than Oracle. Each row names the shape that is served; outside it the statement is refused rather than translated into a different question. | Construct | Served shape — and where it stops | | --- | --- | | ROWNUM | The row cap ports: `WHERE ROWNUM <= n` becomes a limit, a projected `ROWNUM rn` numbers the returned rows from 1, and the nested Oracle pagination idiom (`... WHERE ROWNUM <= 20) WHERE rn > 10`) returns the window it returns in Oracle. Oracle’s counter semantics are preserved rather than approximated: `ROWNUM > 1` matches nothing, as in Oracle, and is never lowered to an offset. Refused (because ROWNUM is assigned before the sort and before grouping, so the obvious lowering would silently change the answer): a cap in the same block as `ORDER BY`, or alongside `DISTINCT` / `GROUP BY` / `HAVING` / an aggregate / a window function; a cap on a level that already has `LIMIT`/`OFFSET`/`FETCH`; `ROWNUM` under an `OR`, compared to a column, or used in GROUP BY / HAVING / ORDER BY / a join condition. For “top n by x” use `ORDER BY x FETCH FIRST n ROWS ONLY`. | | CONNECT BY START WITH | `START WITH ... CONNECT BY [PRIOR] col = col` over a single relation is rewritten to the equivalent recursive query and does the hierarchy walk. Refused by name, never executed with the clause dropped: `LEVEL`, `SYS_CONNECT_BY_PATH`, `CONNECT_BY_ROOT`, `CONNECT_BY_ISLEAF`, `ORDER SIBLINGS BY`, joins or multiple relations, and any relationship that is not a single PRIOR-anchored equality. Carry depth as a column of your own until `LEVEL` lands. | | WITH RECURSIVE | The standard-SQL replacement for `CONNECT BY`, executed to a fixed point. `UNION ALL` only, exactly one recursive CTE per statement, no CTE column list, no aggregate over the recursive relation in the outer projection. Two caps you will meet first: 100 iterations of depth and 1,000,000 accumulated rows; crossing either returns a 400 naming the cap. There is no `CYCLE` clause — carry the visited set as a column. | | SUBSTR | For positive offsets it agrees with Oracle exactly — `SUBSTR('abcdef', 2, 3)` is `'bcd'`. A negative start offset is refused with a 400 naming the divergence rather than answered, because Oracle counts it from the end of the string and these are PostgreSQL semantics. For the last N characters write `SUBSTR(s, LENGTH(s) - (N - 1), N)`. | | INSTR | The 2-argument form ships and is Oracle-equivalent: the 1-based position of the first occurrence, 0 when absent. The 3rd and 4th arguments (start position, occurrence number) are refused rather than evaluated with the extras ignored — a query using them needs restructuring, not a rename. | | DECODE | Ships with the Oracle semantic that trips hand-rewrites: `DECODE` compares NULL-equals-NULL, so a NULL search term matches a NULL input. (If you rewrite to `CASE` instead, remember `WHEN x = NULL` is never true — that arm must become `WHEN x IS NULL`.) Narrower in one way: there is no implicit conversion of search terms to the first term’s datatype, so `DECODE(1, '1', 'y', 'n')` is `'y'` in Oracle and does not match that arm here. Nothing is silently converted or silently matched — the ambiguous case is refused. Keep the input, search terms and results in one type per call. | | SYSDATE SYSTIMESTAMP | Both spellings ship (`SYSDATE` and `SYSDATE()`), folded to an ISO-8601 value at translate time on the same path as `NOW()`, so one statement sees one value. The difference to carry across: the engine answers in UTC, where Oracle’s `SYSDATE` is the database server’s local time. | | TO_DATE | Constant arguments only — both the value and the format model must be quoted literals. Format models served: `'YYYY-MM-DD'`, `'DD-MON-YYYY'`, `'YYYY/MM/DD'`, `'MM/DD/YYYY'`. Where it binds cleanly is the direct comparison, bare column on one side: `WHERE d >= TO_DATE('2026-01-01', 'YYYY-MM-DD')`. Express a date restriction as a range over the bare column rather than through Oracle’s `TRUNC(d) = TO_DATE(...)` idiom, which does not give the same answer here. `TO_CHAR`, `TO_NUMBER` and `TO_TIMESTAMP` do not ship — supply ISO-8601 text. | | (+) outer join | One shape: a single SELECT over exactly two comma-joined tables, every `(+)` marking the same table, each marked predicate a plain `t1.col = t2.col(+)`. It is rewritten to a LEFT OUTER JOIN with the marked table as the optional side, and unmarked conjuncts move to the residual WHERE — which is Oracle-faithful, not a shortcut. Everything else is refused: three or more relations, a mix with explicit JOIN syntax, marks on both sides of one equality, a mark against a constant or inside a function, and `WITH`/`GROUP BY`/`HAVING`/`DISTINCT` alongside it. Refusing matters here because the natural fallback — drop the marks and join — is an INNER join, which silently returns fewer rows. The anti-join `AND t2.id IS NULL` is also refused, the same refusal the hand-written ANSI spelling gets. | | MERGE INTO | One narrow shape: `MERGE INTO t USING (VALUES ...) AS s (c1, ...) ON t.pk = s.k WHEN MATCHED THEN UPDATE SET ... WHEN NOT MATCHED THEN INSERT ...`. It is lowered onto the same upsert plan as `INSERT ... ON CONFLICT`, inheriting its conflict-target rules and its access gate. Refused: a `USING` source that is a table or subquery, conditional `WHEN ... AND`, `WHEN MATCHED THEN DELETE`, `BY SOURCE`/`BY TARGET`, more than one clause of a kind, and an `ON` that is not a conjunction of equalities covering exactly the primary key or one unique index. A set-based MERGE driven by a query has to be restructured. | ## Not supported. Not accepted today. Each row carries the rewrite that gets you the same answer, so you can plan the port rather than discover the gap at cutover. | Construct | What to write instead | | --- | --- | | PIVOT / UNPIVOT | Build the cross-tab with conditional aggregates — `SUM(CASE WHEN k = 'a' THEN v END) AS a, ...` with a `GROUP BY` — which is served over a single table and over a join. What is not expressible is the dynamic form, where the output columns come from the data: the engine plans a fixed projection, so those column names must be known when the statement is written. | | ROWID | Row identity here is the declared primary key, which every registered table has; row version for optimistic concurrency is the value the `If-Match` surface uses. Note that an Oracle `ROWID` is a physical address and is not stable across a reorganisation, so code that persisted one was already relying on something no primary key provides — that logic needs revisiting, not translating. (The row-cap uses of `ROWNUM` are a separate construct and do port — see Preview above.) | | Flashback query | `AS OF TIMESTAMP`, `AS OF SCN` and `VERSIONS BETWEEN` are not accepted, and no session setting makes a SELECT read an earlier state. Historical recovery here is an operator-driven point-in-time restore against a recovery point, which produces a restored database rather than answering a query mid-session — a different tool for a different job. See [Operations](https://originchaindb.com/docs/ops). | | PL/SQL proper | Anonymous `BEGIN` blocks submitted as statements, packages, `%ROWTYPE` / `%TYPE` anchored declarations, `BULK COLLECT` / `FORALL` and autonomous transactions are not accepted. Stored procedures do ship in a different shape — [see below](https://originchaindb.com/docs/compatibility/oracle-dialect#stored-procedures). | | Oracle Net (TNS) | Not planned — a decision rather than a gap. The engine’s SQL wire protocol is PostgreSQL’s, so an Oracle-shaped application connects through a PostgreSQL driver and its SQL is what has to be portable, which is what the rest of this page is about. There is no TNS listener, no Oracle Net negotiation and no OCI surface. Do not read this row as work that is queued. | ## Stored procedures: the boundary. Stored procedures ship in preview, in a PL/pgSQL-shaped subset. This is the part of a port most likely to be mis-scoped, so it is worth being blunt: this is not PL/SQL, and a package body will not move across. What travels is the control flow, the cursors and the exception structure — not the packaging. What ships. - `CREATE PROCEDURE name(p1 TYPE, ...) AS BEGIN ... END`, invoked with `CALL`. A CALL outside a client transaction runs the whole body atomically: an uncaught failure discards every prior statement in the body. - Control flow: `DECLARE`, assignment, `IF`/`ELSIF`/`ELSE`, `CASE`, `LOOP`, `WHILE`, `FOR` over a range and `FOR` over a query, `EXIT`/`CONTINUE`, `RETURN`, `RAISE`, `PERFORM`, dynamic `EXECUTE`, and nested `BEGIN ... END` scopes. - `SELECT ... INTO [STRICT]`, binding the first row’s columns positionally. - Cursors, including a client-returning fetch: `DECLARE c CURSOR FOR` / `OPEN c FOR`, then `FETCH ... INTO` to bind locals, or `FETCH` without `INTO` to deliver rows to the caller. Forward fetch only. - `RETURN QUERY` and `RETURN NEXT` to build a result set for the caller. - `EXCEPTION WHEN ... THEN` handlers on any block, with the block’s own writes rolled back before the handler runs. the boundary, stated plainly - A client CALL cannot collect OUT parameters. Return values to the caller with `RETURN QUERY` or a client-returning `FETCH` instead. CALL arguments are literals. - No packages, no `%ROWTYPE` / `%TYPE` anchored declarations, no `BULK COLLECT` / `FORALL`, and no anonymous blocks submitted as a statement. - No `COMMIT` inside a body, and no autonomous transactions — the body commits as one unit at the end. - Catchable conditions are `OTHERS`, `no_data_found`, `too_many_rows` and `integrity_constraint_violation`. A finer SQLSTATE name is refused at CREATE rather than silently caught as `OTHERS`. - Loops are capped at 100,000 iterations, and a CALL expands to a bounded statement budget — a `RETURN NEXT` loop is not the way to return a large result set; `RETURN QUERY` is. - No loop labels, no `FOREACH ... IN ARRAY`, no dynamic cursors, and no backward fetch directions. A `CALL` issued inside a client `BEGIN`…`COMMIT` joins that transaction rather than opening its own, which forfeits the body’s all-or-nothing property: on a failure the body’s earlier statements stay in your transaction and a later `COMMIT` persists them. Roll back explicitly if that is not what you want. ## Identifier case: the first thing that will bite you. This is not a dialect function, it is a whole-schema property, and it breaks ports before any of the constructs above get a chance to. Oracle folds an unquoted identifier to UPPER case. This engine, like PostgreSQL, folds it to lower case. The consequence is that a ported schema fails in a way that looks like a missing object. Oracle DDL that wrote `CREATE TABLE EMP` stored a table called `EMP`; if you replay the same DDL here it creates `emp`. Then a query that quotes `"EMP"` reports that no such table exists — and you go looking for a failed migration instead of a case rule. per-tenant case resolution is in development A per-tenant identifier case-resolution setting — which would let an instance resolve identifiers the way Oracle does — is in development and is not available on any instance today. Do not plan a migration around it. The workaround below is the supported answer for now, and it is a good habit regardless. What to do today: pick one convention and never leave it. - Preferred — unquoted everywhere. Strip the double quotes from your DDL and from your queries. Every identifier folds to lower case consistently, and Oracle-style unquoted SQL keeps working because it folds too — just in the other direction, to the same place. - Or — quoted everywhere, in one case. If your DDL created `"EMP"`, then every reference to it must also be `"EMP"`, in every query, view and procedure body. This works, but it is unforgiving: one unquoted mention resolves to `emp` and fails. - Never mix. A schema with both `"EMP"` and `emp` is legal and is the worst outcome — two real tables, one typo apart. ``` -- Oracle DDL, replayed as-is. Creates a table named `emp` (folded down). CREATE TABLE hr.EMP (id NUMBER, ename VARCHAR2(30)); -- Fails: there is no "EMP", because the quotes ask for it exactly. SELECT * FROM hr."EMP"; -- error: table not found -- Works: unquoted, folds to `emp`, matches what the DDL created. SELECT * FROM hr.EMP; SELECT * FROM hr.emp; -- the same table ``` ## Worked examples. Real Oracle statements, and what happens to them here. Each block runs unchanged over the HTTP /sql endpoint or through a PostgreSQL driver. ### 1. A sequence-backed insert Both spellings work and share one durable counter, so you can port the Oracle spelling and migrate it later — or not at all. ``` CREATE SEQUENCE hr.emp_seq START WITH 1000 INCREMENT BY 1; -- The Oracle pseudocolumn spelling: ships as-is. INSERT INTO hr.emp (id, ename) VALUES (hr.emp_seq.NEXTVAL, 'KING'); -- The function spelling: the same counter, same durability. INSERT INTO hr.emp (id, ename) VALUES (nextval('hr.emp_seq'), 'CLARK'); -- Read the value the session last took. SELECT hr.emp_seq.CURRVAL FROM dual; -- REFUSED - this is the table-scanning case, and it would mean -- "one sequence value per scanned row", which this surface does not do: -- INSERT INTO hr.emp (id, ename) SELECT hr.emp_seq.NEXTVAL, ename FROM staging.emp; ``` ### 2. NULL handling and FROM DUAL The one difference to remember: an empty string is not NULL here, so `NVL('', 'x')` is `''` rather than Oracle’s `'x'`. ``` SELECT NVL(commission, 0) AS commission, NVL2(manager_id, 'reports', 'top') AS position, DECODE(dept, 10, 'ops', 20, 'sales', 'other') AS dept_name FROM hr.emp; -- FROM DUAL is recognised and stripped: SELECT SYSDATE FROM dual; -- answers in UTC SELECT 1 + 1 FROM sys.dual; -- either spelling, any case ``` ### 3. Row caps and pagination The cap and the nested pagination idiom port. The combination that would silently change meaning is refused, and the refusal names both rewrites. ``` -- Works: a plain cap. SELECT id FROM hr.emp WHERE ROWNUM <= 10; -- Works: the classic nested pagination idiom, rows 11-20. SELECT * FROM ( SELECT a.*, ROWNUM rn FROM ( SELECT id, ename FROM hr.emp ORDER BY id ) a WHERE ROWNUM <= 20 ) WHERE rn > 10; -- Matches NOTHING - exactly as in Oracle, where the counter only -- advances on a returned row. It is not an OFFSET. SELECT id FROM hr.emp WHERE ROWNUM > 1; -- REFUSED: ROWNUM is assigned BEFORE the sort, so this is -- "10 arbitrary rows, then sorted" in Oracle - not the 10 smallest. -- SELECT id FROM hr.emp WHERE ROWNUM <= 10 ORDER BY id; -- Write the top-n you actually meant: SELECT id FROM hr.emp ORDER BY id FETCH FIRST 10 ROWS ONLY; ``` ### 4. Hierarchy and set difference `CONNECT BY` does the walk; `LEVEL` does not ship yet, so carry depth yourself if you need it. ``` -- Walks the reporting tree from the top down. SELECT id, ename FROM hr.emp START WITH manager_id IS NULL CONNECT BY PRIOR id = manager_id; -- The portable spelling of the same walk, with depth carried as a column: WITH RECURSIVE tree AS ( SELECT id, ename, 1 AS depth FROM hr.emp WHERE manager_id IS NULL UNION ALL SELECT e.id, e.ename, t.depth + 1 FROM hr.emp e JOIN tree t ON e.manager_id = t.id ) SELECT id, ename, depth FROM tree; -- MINUS is a spelling of EXCEPT and inherits its behaviour exactly. SELECT region FROM sales.orders MINUS SELECT region FROM sales.returns; ``` ### 5. The (+) outer join The two-table shape is rewritten to a LEFT OUTER JOIN. Anything wider is refused rather than turned into an inner join behind your back. ``` -- Served: two tables, all marks on one side, plain col = col. SELECT e.ename, d.dname FROM hr.emp e, hr.dept d WHERE e.dept_id = d.id(+); -- Which is exactly this, and you may prefer to write it directly: SELECT e.ename, d.dname FROM hr.emp e LEFT OUTER JOIN hr.dept d ON e.dept_id = d.id; ``` ## Next. Everything the base SQL surface serves — joins, aggregates, window functions, CTEs, constraints and transactions — is on the [SQL reference](https://originchaindb.com/docs/sql), and it is what you fall back to wherever an Oracle spelling stops. To get a client connected, start at [Connect a SQL client](https://originchaindb.com/docs/connect-sql-client). If a statement you depend on is refused, the message names the construct — send it to us and it becomes a tracked gap rather than a surprise. Oracle and Java are registered trademarks of Oracle Corporation and/or its affiliates. OriginChainDB is not affiliated with, endorsed by, or sponsored by Oracle Corporation. OriginChainDB implements a compatible subset of the Oracle SQL dialect so that existing SQL can be ported to it; it does not distribute, embed or resell Oracle software, and it does not implement Oracle’s network protocol. --- # The SQL Server wire protocol (TDS) on OriginChainDB Canonical source: https://originchaindb.com/docs/compatibility/sqlserver-tds Sitemap last modified: 2026-09-13T20:18:50.000Z reference · tds # The SQL Server wire protocol OriginChainDB implements TDS, the wire protocol used by Microsoft SQL Server. Point a SQL Server driver at an instance and it connects, authenticates, prepares statements and reads rows against the same store your [HTTP API](https://originchaindb.com/docs/api) calls see. Whether a session may also write depends on its [credential model](https://originchaindb.com/docs/compatibility/sqlserver-tds#credentials): a shared-credential listener reads and writes; a per-user listener is read-only, and refuses every write. This page is the dialect reference: what the protocol carries, what the SQL translator accepts, and — the part people get wrong — what it refuses. read this before you plan a migration The wire is compatible. T-SQL is not the language. A batch that arrives over TDS is parsed with a SQL Server grammar and rewritten, statement by statement, into OriginChainDB's own SQL. That translation covers the query language. It does not cover the procedural one. There are no T-SQL variables, no `@@` functions, no `GO`, no `#temp` tables, and no T-SQL stored procedures. Procedures on OriginChainDB are a PL/pgSQL-shaped subset, which is a different language with different syntax and different error semantics. So a stored procedure written for SQL Server has to be rewritten, not ported. If your application's logic lives in procedures, triggers or scripted batches, budget for that first — it is the largest single cost of moving, and it is not reduced by anything else on this page. [The exact boundary is below.](https://originchaindb.com/docs/compatibility/sqlserver-tds#tsql) preview · off by default · enabled per instance The adapter is built into the shipping engine and is inert until it is configured. No instance gets a TDS listener by default, enabling one is not self-serve, and on production instances today `/v1/capabilities` reports `available: false` for it. Enabling requires an address, a tenant, a credential and TLS material, and it is a preview surface, not a generally available one. Everything described here also runs today over the [SQL endpoint](https://originchaindb.com/docs/sql) and the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client), neither of which needs any of this. [Talk to us](https://originchaindb.com/contact) if you want a listener on your instance. ## Check what is bound, not what is configured The `available` flag under `mssql_tds_dialect` means configuration is complete. It is a precondition for starting a listener, not an observation that one is listening — and, as the next section explains, this adapter deliberately refuses to bind in states where it could not be safe. So check the flag, then the socket, then a real login. ``` # 1. What does this instance say about the adapter? curl -s "https://$OC_HOST/v1/capabilities" | jq '.families.mssql_tds_dialect' # 2. "available" means CONFIGURED, not "a socket is listening". # Check the socket too. nc -vz $OC_SQL_HOST 1433 # 3. Then prove it end to end with a real login, not a port scan. # (sqlcmd, or the JDBC / Go snippets further down this page.) ``` `1433` is the port SQL Server clients expect and the one used throughout this page, but the host and port your instance is published on are shown in the console once a listener exists. Use those. The `limitations` your own instance reports beside `available` are derived from its configuration, and are authoritative for it. ## TLS is mandatory, and it fails closed There is no plaintext path to this listener — not a discouraged one, not a flag, none. That is worth stating plainly because it changes what a misconfiguration looks like: a broken certificate does not produce a degraded connection, it produces no port at all. If nothing answers on `1433`, check the certificate before you check the firewall. It is enforced at four places, each of which fails closed: - At startup. If the certificate and key are missing, empty or unloadable the listener does not bind, and the reason is logged. The check runs before the bind, so no port is ever presented as up. - At PRELOGIN. A client that offers no encryption, or asks for it to be off, is answered with the protocol byte meaning encryption is required by the server and the connection is closed cleanly, before the login packet is read. It is not downgraded and not accepted-then-ignored. - Against login-only encryption. Encrypting the login and then reverting to plaintext — which a real SQL Server will do — is rejected by design, not merely unimplemented. It protects the credential while streaming every query and every result row in clear, which for a database leaks exactly what the credential would have unlocked. - In the type system. The command loop takes an encrypted stream that only a completed handshake can produce, so "no session runs unencrypted" is a compile-time property of the adapter rather than a rule someone has to remember. Set `encrypt=true` and leave `trustServerCertificate` at `false` so your client actually validates the chain. The handshake itself is the TDS-framed variant that commercial SQL Server drivers require during negotiation; the stream switches to ordinary TLS records from the login packet onward. This is the opposite of OriginChainDB's MySQL-wire adapter, which fails open — it binds without a certificate and then requires nothing. The two adapters behave differently on purpose, and a habit learned on one is the wrong habit on the other. ## Connecting with a real client Only drivers that have actually been driven against the engine appear here, at the versions they were driven at. ### Java — mssql-jdbc 13.4.0 and 12.10.1 ``` // Measured against mssql-jdbc 13.4.0 and 12.10.1 - 34 of 34 checks passed. String url = "jdbc:sqlserver://" + System.getenv("OC_SQL_HOST") + ":1433" + ";databaseName=" + System.getenv("OC_DATABASE") + ";encrypt=true" // mandatory - there is no other way in + ";trustServerCertificate=false" + ";hostNameInCertificate=" + System.getenv("OC_SQL_HOST") + ";loginTimeout=30"; try (Connection conn = DriverManager.getConnection( url, System.getenv("OC_DB_USER"), System.getenv("OC_DB_PASSWORD"))) { try (PreparedStatement ps = conn.prepareStatement( "SELECT TOP 50 id, customer_ref, amount_cents" + " FROM invoices WHERE status = ? ORDER BY amount_cents DESC")) { ps.setString(1, "open"); try (ResultSet rs = ps.executeQuery()) { while (rs.next()) { System.out.println(rs.getInt("id") + " " + rs.getString("customer_ref")); } } } } ``` ### Go — microsoft/go-mssqldb 1.11.0 ``` // Measured against microsoft/go-mssqldb 1.11.0 - 21 of 21 checks passed. package main import ( "database/sql" "fmt" "net/url" "os" _ "github.com/microsoft/go-mssqldb" ) func main() { q := url.Values{} q.Set("database", os.Getenv("OC_DATABASE")) q.Set("encrypt", "true") // mandatory q.Set("TrustServerCertificate", "false") // validate the chain dsn := url.URL{ Scheme: "sqlserver", User: url.UserPassword(os.Getenv("OC_DB_USER"), os.Getenv("OC_DB_PASSWORD")), Host: os.Getenv("OC_SQL_HOST") + ":1433", RawQuery: q.Encode(), } db, err := sql.Open("sqlserver", dsn.String()) if err != nil { panic(err) } defer db.Close() rows, err := db.Query( "SELECT TOP 50 id, customer_ref FROM invoices WHERE status = @p1", "open") if err != nil { panic(err) } defer rows.Close() for rows.Next() { var id int var ref string if err := rows.Scan(&id, &ref); err != nil { panic(err) } fmt.Println(id, ref) } } ``` Worked, page-sized versions of both, plus the prepared-statement path, are in the [TDS examples](https://originchaindb.com/docs/examples/sqlserver). ## Credential models A listener runs in one of two models, and the choice is not yours to make freely — it follows from whether the tenant has any database authorization to enforce: a database RBAC grant or a row-level security policy. Either one is enough; a row-security policy is not the lighter case. | Model | What the login carries | Applies when | | --- | --- | --- | | shared | One listener-wide credential. Every session arrives as the same identity. | The tenant has no database RBAC grants and no row-level security policies at all. A shared credential cannot enforce a per-user grant or evaluate a per-user row policy, so the engine will not let the two be combined. | | per-user | Each database user logs in with its own password, under the mandatory TLS, and its own grants apply. | The tenant has any database RBAC grant, or any row-level security policy. This is the only model that can serve it, and the session is read-only. | a per-user listener is read-only A per-user TDS session refuses every write. Not only the writes a grant disallows — the listener marks the session read-only the moment it admits an authenticated identity, before any grant is consulted, because there is no write-parity bridge on this adapter. `INSERT`, `UPDATE`, `DELETE` and DDL all come back refused, and no configuration turns this on. Put that beside the table above and the consequence is the one to plan around: a tenant with a single RBAC grant or a single row-level security policy cannot write over TDS at all. That tenant must be on per-user login, and per-user login does not write. No combination of settings yields an authorizing, writing TDS session today. So every write shown on this page and under [TDS examples](https://originchaindb.com/docs/examples/sqlserver) — including statements in the translated set, which translate correctly and are still refused on a per-user session — applies to a shared-credential listener only. If your tenant has any authorization to enforce, write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client), and keep the TDS listener for reads. Put those two rows together and there is a real state that neither model can serve: a tenant that carries database RBAC or a row-level security policy but has no database user with a password verifier. Shared login is inadmissible because that authorization is present; per-user login has no account to admit, so the listener refuses to bind — by design, rather than binding something it could not enforce. This is not an exotic corner: it is the ordinary state of a tenant whose access has only ever been through API tokens. Creating one enabled database user with a password is the whole fix. ## The T-SQL boundary, exactly A batch is parsed with a SQL Server grammar, the syntax tree is rewritten into OriginChainDB's SQL, and each statement is re-rendered and executed through the same engine every other surface uses. The rule the translator follows is soundness over coverage: anything without an exact equivalent is refused with a clear error and nothing executes. It never guesses, and it never returns a plausible wrong answer. That makes the boundary easy to test and unpleasant to discover late. Here it is. ### Translated | T-SQL you write | What it becomes | | --- | --- | | [bracket] identifiers | A plain name unwraps; a name with spaces or punctuation, or a reserved word such as `[order]`, becomes a double-quoted identifier. Every position a statement can put one in. | | SELECT TOP n · TOP (n) | `LIMIT n` | | OFFSET n ROWS FETCH NEXT m ROWS ONLY | `LIMIT m OFFSET n`. Always lowered — a `FETCH` left alone would be ignored and silently return every row. | | ISNULL · LEN · CHARINDEX · IIF | `COALESCE`, `LENGTH`, `POSITION`, `CASE` | | GETDATE · SYSDATETIME · SYSUTCDATETIME | One statement-time UTC value, folded so every call in a statement agrees. | | CONVERT(type, expr) | `CAST(expr AS type)`. The UTF-16 character types (`NVARCHAR`, `NCHAR`, `NTEXT`) collapse onto the engine's string types and `UNIQUEIDENTIFIER` onto its UUID type, in casts and in DDL column definitions alike. Only mappings that change no value representation are made: `BIT`, `MONEY` and `DATETIME2` are deliberately left alone, so the engine refuses them with a clear error rather than accept mis-typed data. | | a + b (string concatenation) | `CONCAT(a, b)`, but only when both operands are provably string. Two numbers stay arithmetic; one string and one unknown is ambiguous in T-SQL itself and is refused. | | SUBSTRING · REPLACE · LTRIM · RTRIM · UPPER · LOWER | Pass through unchanged — they are engine natives. | ``` -- Every line below is rewritten and executed. Send it as an ordinary -- SQL_BATCH, or as the statement text of a parameterized RPC. SELECT TOP 20 [order].[id], [customer name] FROM [order] WHERE ISNULL([status], 'open') = 'open' ORDER BY [amount cents] DESC; SELECT id, LEN(customer_ref) AS ref_len, CHARINDEX('-', customer_ref) AS dash_at, IIF(amount_cents > 100000, 'large', 'small') AS bucket, CONVERT(NVARCHAR(32), amount_cents) AS amount_text FROM invoices ORDER BY id OFFSET 40 ROWS FETCH NEXT 20 ROWS ONLY; ``` ### Refused, cleanly Each of these comes back as an error token naming the construct and, where one exists, the portable rewrite. Nothing runs, nothing is half-applied, and the connection survives. | Construct | Why, and what to write instead | | --- | --- | | @var · @@FUNCTION | Batch-level variables and built-in globals are not modelled. Bind a parameter, or compute the value in your application. | | #temp objects | Per-session temporary objects are not modelled. Create an ordinary table. | | GO | A client-tool batch separator, not a server statement. Send one statement per batch. It is caught before parsing, because the grammar would otherwise read a trailing `GO` as a column alias. | | MERGE | Use `INSERT ... ON CONFLICT`. | | SELECT ... INTO newtable | The engine parses the `INTO` clause and drops it, so passing it through would report a successful table copy that created nothing. Use `CREATE TABLE` then `INSERT ... SELECT`. | | DATEADD · DATEDIFF | The substrate has no calendar timestamp type — timestamps are ISO-8601 strings — so there is no sound mapping. Do date arithmetic in your application. | | CONVERT with a style code | Style codes and `USING` charsets are formatting semantics that are not reproduced. Format in your application. | | TOP ... PERCENT · WITH TIES | No `LIMIT` equivalent. The same applies to `FETCH ... PERCENT / WITH TIES`. | | TOP in UPDATE / DELETE | Constrain the statement with a WHERE predicate. | | Anything the SQL Server grammar cannot parse | Reported with the parse failure rather than attempted. A final safety net also refuses any translated output still carrying an un-rewritten bracket identifier, rather than letting it reach the engine as a wrong name. | ``` -- Each of these is REFUSED with an explicit error. Nothing executes, -- nothing is half-applied, and the connection stays open. DECLARE @cutoff INT = 100000; -- @variables are not modelled SELECT @@VERSION; -- @@functions are not modelled SELECT * INTO #recent FROM invoices; -- # temp objects are per-session GO -- a client-tool directive, not a statement MERGE invoices AS t USING staged AS s -- use INSERT ... ON CONFLICT instead ON t.id = s.id WHEN MATCHED THEN UPDATE SET t.status = s.status; SELECT DATEADD(day, -7, GETDATE()); -- no calendar timestamp type SELECT CONVERT(VARCHAR(10), created_at, 112); -- style codes are formatting SELECT TOP 10 PERCENT * FROM invoices; -- no LIMIT equivalent DELETE TOP (100) FROM invoices; -- constrain with WHERE instead SELECT * INTO archive FROM invoices; -- SELECT ... INTO would create nothing ``` ### Three approximations to know about These translate, and for ordinary inputs they agree with SQL Server. They are listed because for unusual inputs they do not: - `LEN` does not reproduce T-SQL's trailing-space trim, so `LEN('a ')` counts the spaces. - `GETDATE()` yields an ISO-8601 UTC string, because there is no datetime type to yield instead. - `CONCAT` treats NULL as empty where T-SQL's `+` propagates it. Only reachable where both operands are provably string, which is the only case that translates at all. ## Stored procedures: the distinction that costs money TDS carries a remote-procedure-call path, and OriginChainDB serves the part of it every driver depends on: `sp_executesql`, `sp_prepare`, `sp_execute`, `sp_prepexec` and `sp_unprepare`. Seeing those names work is what makes people assume stored procedures work. They are not your procedures. That family is the protocol's own machinery — it is how a driver ships a parameterized query and reuses its plan handle. Serving it is what makes a `PreparedStatement` work on its second execution and what makes ODBC work on its first. It says nothing about running procedures you wrote. Calling a user stored procedure over the wire is refused, and so is every other procedure, by numeric id or by name. Also refused on this path, cleanly rather than mis-served: the server-side cursor family, OUTPUT value parameters (there is no variable-assignment surface that could produce one), bulk load, and multiple active result sets. Temporal parameter values decode best-effort in this preview — integer, float, bit, string, numeric and unique-identifier parameters are exact. OriginChainDB does have procedures. They are written in a PL/pgSQL-shaped subset, created and called over the [SQL endpoint](https://originchaindb.com/docs/sql) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client), and they are a genuinely different language. Here is what a small one looks like on both sides — note that nothing about the T-SQL version survives except its intent: ``` -- SQL Server (T-SQL). This does NOT run on OriginChainDB. CREATE PROCEDURE dbo.settle_invoice @id INT, @settled INT OUTPUT AS BEGIN DECLARE @amount INT; SELECT @amount = amount_cents FROM invoices WHERE id = @id; IF @amount IS NULL THROW 50001, 'no such invoice', 1; UPDATE invoices SET status = 'settled' WHERE id = @id; SET @settled = @@ROWCOUNT; END; -- OriginChainDB. A different procedural language, so this is a rewrite - -- not a port. Issue it over the SQL endpoint or the PostgreSQL wire. CREATE PROCEDURE settle_invoice(IN p_id INT, INOUT p_settled INT) LANGUAGE plpgsql AS $$ DECLARE v_amount INT; BEGIN SELECT amount_cents INTO v_amount FROM invoices WHERE id = p_id; IF v_amount IS NULL THEN RAISE EXCEPTION 'no such invoice'; END IF; UPDATE invoices SET status = 'settled' WHERE id = p_id; GET DIAGNOSTICS p_settled = ROW_COUNT; END; $$; ``` For anything larger, estimate the rewrite the way you would estimate a rewrite: by counting the procedures and reading them, not by assuming a compatibility layer will absorb them. ## How far this has been proven Interop is measured by driving Microsoft's own shipped drivers against the real accept loop — not a client written against our own encoder, which is how earlier claims about this adapter turned out to be wrong while every in-house test was green. | Driver | Version | Result | | --- | --- | --- | | mssql-jdbc | 13.4.0 | 34 / 34 | | mssql-jdbc | 12.10.1 | 34 / 34 | | microsoft/go-mssqldb | 1.11.0 | 21 / 21 | Thirteen consecutive rounds, all green, measured on 2026-09-07. What those rounds cover: connecting and logging in over the framed TLS handshake; DDL; insert; select with rows decoded and compared; the whole prepare / execute / unprepare family asserted through the driver's own prepared-statement handle rather than through our encoder; batch execution; and a wrong credential being refused. Both drivers are required — each one exercises framing, flush and row-count behaviour the other tolerates, and a single-driver run has already been shown to miss half of a defect set. verified, not continuously verified That battery runs in no automated job. It was executed deliberately and it passed, so the accurate word is verified — as of that date, against those driver versions, against that engine build. A regression introduced by a later build would not be caught by a scheduled run, because there is not one. Treat the figures above as a measurement with a date on it, and re-run the battery against the build you intend to deploy rather than assuming they still hold. ## Limits, in one place - The T-SQL procedural language is not implemented. No variables, no built-in globals, no control-flow batches, no `GO`, no temp tables, no T-SQL stored procedures, triggers or functions. Procedures are a PL/pgSQL-shaped subset and are a rewrite. - This is a preview surface. It is not generally available, it is off on every instance unless someone turned it on, and it is not self-serve. - TLS is required absolutely. No certificate means no listener, and a client that will not encrypt is refused before it can log in. - Server-side cursors, bulk load and multiple active result sets are refused — cleanly, but refused. So is calling a user stored procedure, and so are OUTPUT value parameters. - Date and time arithmetic does not translate. There is no calendar timestamp type behind this surface. - Interop is verified, not continuously verified, against the three driver versions named above and no others. Other TDS clients may work; they have not been measured, so we do not claim them. If any of that is disqualifying, the same data is reachable today over the [SQL endpoint](https://originchaindb.com/docs/sql) and the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client), both of which are ordinary supported surfaces with none of the caveats on this page. Every third-party name on this page is used to describe protocol or dialect compatibility, and for no other purpose. Microsoft, SQL Server and T-SQL are trademarks of the Microsoft group of companies. PostgreSQL is a registered trademark of the PostgreSQL Community Association of Canada. MySQL is a registered trademark of Oracle Corporation and/or its affiliates. mssql-jdbc and go-mssqldb are Microsoft's own drivers and are named only to identify the versions this compatibility was measured against. OriginChainDB is not affiliated with, endorsed by, or sponsored by any of them, and does not distribute their software. OriginChainDB is not SQL Server: it is an independent database that implements the TDS wire protocol so that existing SQL Server clients can connect to it, and that accepts part of the T-SQL query language. --- # Connect an Elasticsearch client - _bulk, Query DSL, aggs Canonical source: https://originchaindb.com/docs/connect-elasticsearch Sitemap last modified: 2026-09-23T08:33:23.000Z reference · elasticsearch # Connect an Elasticsearch client Connect an Elasticsearch client to OriginChainDB to index documents, run supported Query DSL searches, and aggregate results. available on every instance There is nothing to switch on. Every instance answers the Elasticsearch API at https:///v1/tenants//es — that whole path is what a client takes as its node URL, with your API key as the bearer token. ## What you can connect. Any client that speaks the Elasticsearch 7.x REST API. We verify against the official @elastic/elasticsearch Node client — it connects, clears its product check, and runs a full index → search → aggregate → delete lifecycle unchanged. Dashboards and app code that issue the Query DSL keep working; you change the endpoint, not the queries. The cluster reports version 7.14.2, so pin your client to the 7.x line. ## Connection details. Point the client at your endpoint and authenticate with an API key from the console. info() is the handshake — it is what the client uses to confirm it is talking to Elasticsearch. ``` const { Client } = require('@elastic/elasticsearch') const es = new Client({ node: 'https://', // your instance URL + /v1/tenants//es auth: { apiKey: '' } }) await es.info() // clears the product check; reports version 7.14.2 ``` ## Insert data. Index a single document by _id. It is searchable the instant the call returns — there is no refresh interval to wait on. ``` await es.index({ index: 'shop.products', id: 'sku-8842', document: { name: 'Carbon Marathon', brand: 'Aero', price: 149 } }) ``` Load many at once with _bulk (newline-delimited actions). This is the fast path for backfills and ingest. ``` await es.bulk({ operations: [ { index: { _index: 'shop.products', _id: 'sku-1207' } }, { name: 'Trail 24', brand: 'Aero', price: 89 }, { index: { _index: 'shop.products', _id: 'sku-3355' } }, { name: 'City Runner', brand: 'Metro', price: 72 } ]}) ``` Or straight over HTTP — the same NDJSON body a stock Elasticsearch cluster takes: ``` POST /_bulk {"index":{"_index":"shop.products","_id":"sku-1207"}} {"name":"Trail 24","brand":"Aero","price":89} ``` _update is a partial merge — the fields you send are updated and every other field is preserved. ``` await es.update({ index: 'shop.products', id: 'sku-8842', doc: { price: 139 } // name, brand, ... untouched }) ``` ## Search and aggregate. The Query DSL you already write — bool, match, term, filters — with aggregations in the same request. ``` await es.search({ index: 'shop.products', query: { bool: { must: [{ match: { name: 'marathon' } }], filter: [{ term: { brand: 'Aero' } }] } }, aggs: { by_brand: { terms: { field: 'brand' } } } }) ``` Also supported: _count, _msearch, _mget, search_after deep paging, collapse, _delete_by_query and _update_by_query (both honor max_docs), and _reindex into a fresh index. Read a document straight back with GET or HEAD /:index/_doc/:id, or /:index/_source/:id for the bare document; ask what a field looks like with _field_caps; and guard a bulk write with if_seq_no — a stale precondition comes back as a per-item 409 and every other item in the batch still lands. ## Security comes with the search. Row-level security and column masking apply to the search itself, not just to the documents it returns. A caller who cannot see a row will not find it through a query, and a query that searches a masked column is refused rather than answered — the match set can't be used to reconstruct a value the mask hides. This is enforcement a bolt-on search cluster can't give you, because it never sees your database's policies. ## Limits, stated plainly. This is a drop-in for the API a real client and its dashboards exercise — not the entire Elasticsearch surface. What we don't answer yet fails closed with an explicit error; it never returns a wrong or silently-partial result. - Aggregations: terms, metrics, range, histogram, date_histogram, percentiles, top_hits, sub-aggregations and pipeline aggs are in, and so is filters — named buckets, each its own query, with sub-aggregations over exactly that bucket. Send filters in its own request: it is answered with one search per filter. extended_stats, composite and nested are not answered yet, and missing on a metric aggregation is refused rather than filled in with a value you did not choose. - Suggesters: the term suggester corrects a word against what is actually indexed in a field, with suggest_mode honored. Options carry the corrected text and a score, and no document frequency — that number is not available to this API, and we would rather omit a field than report one we cannot stand behind. The phrase and completion suggesters are refused by name. - Scoring: function_score and per-field boosts are in; script_score and rescore are refused. - Deep paging: search_after within a 10,000-hit window; Point-in-Time (PIT) is not offered. - Cluster management: ILM, snapshots, and index templates are managed by OriginChainDB, not over the ES API. Vector / kNN search lives on a [dedicated surface](https://originchaindb.com/docs/vector). - Version: the cluster reports 7.14.2; pin clients to the 7.x line. Feature coverage is a separate question from scale — see [Full-text search](https://originchaindb.com/docs/fts) for how the index is built and what a single node holds. Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Connect a SQL client - psql, DBeaver, pgAdmin, ODBC Canonical source: https://originchaindb.com/docs/connect-sql-client Sitemap last modified: 2026-09-24T16:37:44.000Z reference · sql clients # Connect a SQL client Connect a PostgreSQL-compatible client to the OriginChainDB SQL endpoint and check which SQL features it supports. limited preview · rolling out Wire-protocol access is in limited preview and is being enabled gradually. Self-serve enablement covers single-node instances for now - for HA or multi-node configurations, talk to us. If your instance doesn't show a SQL access panel in the console yet, the listener isn't switched on for you - everything on this page also works over the [SQL endpoint of the HTTP API](https://originchaindb.com/docs/sql) today. ## What you can connect. Any client or driver that speaks the PostgreSQL protocol can connect. The three we verify against are psql, DBeaver, and pgAdmin - walkthroughs for each are below. Language drivers that use the same protocol (libpq, JDBC's PostgreSQL driver, node-postgres, psycopg, pgx, and friends) connect the same way, and so does the standard PostgreSQL ODBC driver - [DSN details below](https://originchaindb.com/docs/connect-sql-client#odbc). not supported MySQL Workbench and SQL Server Management Studio (SSMS) cannot connect. They speak different wire protocols (the MySQL protocol and TDS respectively), not PostgreSQL's - no connection setting will make them work. Don't spend an afternoon on it; use one of the three clients above. ## Connection details. Host, port, username and password are per-database and live in the [console](https://app.originchaindb.com/): open your instance and look for the SQL access panel. The port is PostgreSQL's stock 5432; authentication is SCRAM-SHA-256 (the flow every current Postgres client speaks), and the connection must use TLS - keep `sslmode=require`. Your existing Postgres client speaks the wire protocol straight to the instance — the same one your [HTTP API](https://originchaindb.com/docs/api), vectors and full-text indexes live in. No driver, no ORM, no schema changes. ``` host # console -> your instance -> SQL access port 5432 # PostgreSQL's stock port - the panel shows it too database postgres # cosmetic - see the note below this block user # shown in the SQL access panel password # shown in the SQL access panel sslmode require # Or as one connection URI: postgresql://:@:5432/postgres?sslmode=require ``` the database field is cosmetic Which database you connect to is decided by your credentials, not by the `database` field. The server reports a single database named `postgres` to every client, whatever you type there. Leave it as `postgres` so GUI clients resolve their navigation tree cleanly - but know that typing something else neither connects you elsewhere nor protects anything. (The console's connection string may name a different database - same cosmetic field, connects identically.) One more gate before the first connect: the SQL port is reachable only from addresses on your instance's IP Access List - it is never open to the world. An empty list means the port stays closed and every client times out. Add the address you connect from (console → your instance → Network access) and retry. ## Client walkthroughs. Each tab is the short version: the fields that matter and the one gotcha worth knowing in advance. pgAdmin gets a fuller step-by-step just below, and ODBC consumers have [a section of their own](https://originchaindb.com/docs/connect-sql-client#odbc). DBeaver — Connect to a database MainPostgreSQLSSL Server Host Port 5432 Database postgres Authentication Database Native Username Password •••••••••••• ✓Save ConnectedTest ConnectionFinish DBeaver → Database → New Database Connection → PostgreSQL. Fill the Main tab as above; on the SSL tab tick Use SSL and set SSL mode to require. All values come from the console’s SQL access panel. #### psql ``` # One line - paste the values from the console's SQL access panel: psql "host= port=5432 dbname=postgres \ user= password= sslmode=require" # Gotcha: \copy (and the COPY protocol generally) is not supported yet. # Bulk-load with a multi-row INSERT ... VALUES (...), (...) instead, # or use the HTTP ingest API. Everything else psql does day-to-day - # queries, \d, \dt, \df, PREPARE/EXECUTE, transactions - works. ``` #### DBeaver ``` 1. New Database Connection -> PostgreSQL (the stock PostgreSQL driver - no custom driver needed). 2. Host, Port, Username, Password: paste from the console's SQL access panel. Database: leave "postgres". 3. SSL tab -> enable SSL. Test Connection -> Finish. Gotcha: the database navigator shows exactly ONE database, named "postgres", regardless of what your database is called in the console. Your table namespaces appear as schemas under it (the default namespace shows as "public"). Don't enable "Show all databases" and go hunting for a database with your name on it; there isn't one. ``` #### pgAdmin ``` 1. Add New Server -> General: any name you like. 2. Connection tab: Host name/address, Port, Username, Password from the console's SQL access panel. Maintenance database: leave "postgres". 3. Parameters tab: SSL mode = Require. Save. Gotcha: pgAdmin's Dashboard panels read PostgreSQL's statistics tables (pg_stat_activity and friends), which this adapter does not serve - expect empty panels or a polite error there. The object browser and the Query Tool are the supported surfaces, and both work. ``` ## pgAdmin 4, step by step. The tab above is the short version; this is the full register-server pass. Open the console's SQL access tab first - every value below comes from it, and the password is shown once, when it's generated. Copy it before you leave the page. Register — Server× GeneralConnectionParametersSSL Host name/address Port 5432 Maintenance database postgres Username Password •••••••••••• Save password? CloseResetSave pgAdmin 4 → right-click Servers → Register → Server. On the Connection tab enter the values above; on the SSL tab set SSL mode to Require. Everything comes from the console’s SQL access panel — copy the password when it is shown. 1. Register the server. Right-click Servers in the object explorer → Register → Server. On the General tab, give it any name you like. 2. Connection tab. Host name/address: your database host from the SQL access tab - it looks like `..db.originchain.ai`. Port: `5432`. Maintenance database: `originchain`, as the panel shows it - the field is cosmetic (see the note under connection details), and the browser tree shows a single database named `postgres` either way. Username and Password: paste from the panel; tick Save password if you want pgAdmin to keep it. 3. SSL mode = Require. On the Parameters tab in current pgAdmin releases; older releases put it on a dedicated SSL tab next to Connection. There is no authentication setting to pick - pgAdmin negotiates SCRAM-SHA-256 automatically. 4. Save. The object browser and the Query Tool are the supported surfaces. The Dashboard panels read PostgreSQL's statistics tables (`pg_stat_activity` and friends), which this adapter does not serve - expect empty panels there, not data. connection timed out? That's the IP Access List, not pgAdmin. The SQL port is reachable only from addresses on your instance's list, and an empty list means the port stays closed - every connection attempt times out. Add the address you connect from (console → your instance → Network access) and retry. Sessions running under a database-user identity - which is what the SQL access tab issues - are read-only: every query is filtered by row-level security and column masking for that user, exactly as the HTTP API, and write statements are refused with SQLSTATE `0A000`. Writes go through the [HTTP SQL endpoint](https://originchaindb.com/docs/sql) - the full story is under the limits section below. ## ODBC (psqlODBC). The standard PostgreSQL ODBC driver (psqlODBC) works - it speaks the same wire protocol as everything above. Install it, pick the Unicode variant, and map the DSN fields to the same values from the console's SQL access tab: ``` Driver PostgreSQL Unicode # psqlODBC - the stock PostgreSQL ODBC driver Server # looks like ..db.originchain.ai Port 5432 Database originchain # cosmetic - see the note under connection details User Name # console -> your instance -> SQL access Password # shown once, when it's generated - copy it then SSL Mode require ``` Or as one connection string, for anything that takes a DSN-less connection: ``` Driver={PostgreSQL Unicode};Server=.ap-south-1.db.originchain.ai;Port=5432;Database=originchain;Uid=;Pwd=;SSLmode=require; ``` excel, power bi, and friends: read-only ODBC consumers connect under the database-user credential from the SQL access tab, and those sessions are read-only. Refreshing a workbook or report works, and every query is filtered by row-level security and column masking for that user - a report shows exactly what its user is allowed to see. Anything that tries to write back over the connection is refused with SQLSTATE `0A000`; writes go through the [HTTP SQL endpoint](https://originchaindb.com/docs/sql). First connection timing out? Same cause as pgAdmin above - the IP Access List. Add your address under Network access first. ## SQL compatibility. The honest matrix. Works over wire means your SQL client runs it directly (and the [HTTP SQL endpoint](https://originchaindb.com/docs/sql) runs it too). API only means the capability exists, but through the HTTP API rather than your SQL client. Not yet means exactly that - we'd rather tell you here than have you find out in production. One caveat spans every row: write statements over the wire apply to a single-node instance without row or column controls. A session running under a database-user identity, and any replicated configuration, gets a read-only wire - see the limits below. | Capability | Status | Notes | | --- | --- | --- | | SELECT | works over wire | Simple and extended protocol, text and binary result formats. Typed columns come back with their real PostgreSQL types (NUMERIC, TIMESTAMPTZ, DATE, UUID, BYTEA, JSON, 1-D arrays, ...). | | INSERT / UPDATE / DELETE | works over wire | Including `RETURNING` and `INSERT ... ON CONFLICT`. Constraint failures surface real SQLSTATEs (23505, 23502, 23503, ...), so driver error handling works unchanged. | | transactions | works over wire | `BEGIN` / `COMMIT` / `ROLLBACK`, plus real savepoints (`SAVEPOINT` / `ROLLBACK TO` / `RELEASE`). Isolation is READ COMMITTED - see the limits below for what happens if you ask for more. | | prepared statements | works over wire | Both kinds: the protocol-level flow every driver uses (Parse/Bind/Execute, text and binary parameters for the common scalar types) and SQL-level `PREPARE` / `EXECUTE` / `DEALLOCATE`. | | stored procedures | works over wire | `CREATE PROCEDURE` / `CREATE FUNCTION`, `CALL` (including row-returning bodies), `DROP PROCEDURE / FUNCTION`. Routines are shared with the HTTP API - create on one surface, call from the other. | | views | API only | `CREATE VIEW` and queries against views run over the [HTTP SQL endpoint](https://originchaindb.com/docs/sql). Over the wire, `CREATE VIEW` is refused and a view's name doesn't resolve yet. | | materialized views | API only | Plain `CREATE MATERIALIZED VIEW name AS select` works over the [HTTP SQL endpoint](https://originchaindb.com/docs/sql) and installs a real materialized view (`OR REPLACE` and `WITH NO DATA` are refused); [the dedicated flow](https://originchaindb.com/docs/schemas/materialized-views) drives the same engine. Over the wire it is refused - see the limits below. | | triggers | not yet | No trigger surface exists on either the wire or the HTTP API. `CREATE TRIGGER` is refused everywhere; there is no partial or emulated version to be surprised by. | DDL over the wire, for completeness: `CREATE TABLE` and `DROP TABLE` are real (the drop cascades through rows and indexes, same as the API path). `CREATE INDEX`, `CREATE SEQUENCE` and `ALTER TABLE` are API-only for now. `WITH` (CTEs) is refused over the wire. `COPY` (both directions) is not supported yet on any client. ## Limits, stated plainly. ### SERIALIZABLE is refused, not faked. Wire connections run at READ COMMITTED. If you ask for more - `SET TRANSACTION ISOLATION LEVEL SERIALIZABLE`, `BEGIN ISOLATION LEVEL REPEATABLE READ`, or the equivalent session default - the server returns an error (SQLSTATE `0A000`) naming the level it can't deliver. It never acknowledges the request and quietly gives you READ COMMITTED anyway. That's deliberate: a database that says yes and delivers less is lying to your correctness logic - write-skew protection you believe you have but don't is worse than an error at connection time. Levels at or below READ COMMITTED are accepted as-is, so drivers that set read-committed on connect work unchanged. Serializable transactions are available - through the HTTP transactions API, on single-node configurations. ### Materialized-view DDL lives on the HTTP endpoint, not the wire. Plain `CREATE MATERIALIZED VIEW name AS select` works over the [HTTP SQL endpoint](https://originchaindb.com/docs/sql) - it is populated at creation, and lands in the same refresh and inspection flow as views defined through [the materialized views reference](https://originchaindb.com/docs/schemas/materialized-views). Two forms are refused rather than half-honoured: `OR REPLACE` (drop, then create) and `WITH NO DATA`. Over the wire, `CREATE MATERIALIZED VIEW` is refused. ### Row and column controls come with an identity - and a read-only wire. Where column masks or row-security policies exist, the wire never serves data without knowing who is asking: enabling SQL access on an instance that carries them requires database-user credentials, and a listener that cannot resolve one refuses to serve those objects rather than read them in the clear. When a session does run under a database-user identity, every `SELECT` is filtered and masked for that user, per session, exactly as the HTTP API - and the session is read-only: every write verb (`INSERT` / `UPDATE` / `DELETE`, `CREATE` / `DROP` / `ALTER` / `TRUNCATE`, `CALL`, `GRANT` / `REVOKE`) is refused with SQLSTATE `0A000` and points you at the [HTTP SQL endpoint](https://originchaindb.com/docs/sql) for the write. Replicated configurations get the same read-only wire regardless of identity - an inline wire write would bypass the replication path. ### Multi-node configurations refuse what they can't guarantee. The same honesty rule applies across configurations: on a [sharded or replicated setup](https://originchaindb.com/docs/multi-node), surfaces that can't uphold their guarantee are refused with a clear error rather than served with silently weaker semantics. Concretely: serializable API transactions are single-node only - a sharded configuration refuses to open one - and on replicated setups they are further restricted to point reads. If a request succeeds, its stated guarantee holds; if we can't hold it, you get an error you can code against. --- # Connect over the MySQL and SQL Server wire protocols Canonical source: https://originchaindb.com/docs/connect-wire-protocols Sitemap last modified: 2026-09-23T08:33:23.000Z reference · wire protocols # MySQL and SQL Server wire protocols Connect MySQL or SQL Server clients to an enabled wire listener and use the supported protocol and SQL features. off by default · enabled per tenant Both adapters are built into the engine and are inert until configured. No tenant gets a MySQL or TDS listener by default, and enabling one is not self-serve today - [talk to us](https://originchaindb.com/contact) and we will tell you whether your tenant can be served at all, which the [credential models](https://originchaindb.com/docs/connect-wire-protocols#credentials) section explains. Everything described here also works today over the [SQL endpoint of the HTTP API](https://originchaindb.com/docs/sql) and over the PostgreSQL wire. ## Check what is bound, not what is configured The `available` flag in `/v1/capabilities` means configuration is complete. It is a spawn precondition, not an observed bound port. A listener can report `available: true` and still have refused to bind - which is exactly what the two refusals below do. So check the socket, then check a real login. ``` # 1. Does the tenant report the listener as CONFIGURED? # The families sit at the TOP level of the payload, not under a wrapper. curl -s "https://$OC_HOST/v1/capabilities" \ | jq '.mysql_wire_dialect.available, .mssql_tds_dialect.available' # 2. Is a socket actually LISTENING? # (1) can say "configured" while (2) refused to bind, so run both - # and give the port scan a control, or a silent "no" proves nothing. nc -vz $OC_SQL_HOST 443 # control: this MUST answer nc -vz $OC_SQL_HOST 3306 # MySQL wire nc -vz $OC_SQL_HOST 1433 # TDS # If the control line fails too, your path is blocked and you have # measured your own network, not our listener. # 3. Prove it end to end with a real login, not a port scan. mysql --host=$OC_SQL_HOST --port=3306 --user=$OC_DB_USER \ --ssl-mode=REQUIRED -e "SELECT 1" ``` There is no default port. Each listener binds whatever address it is configured with, and we follow the convention each client expects - `3306` for the MySQL wire, `1433` for TDS - so a client's own defaults usually need no override. The host and port your tenant is actually published on are the ones we give you when the listener is enabled. Use those, not the examples on this page. ## Credential models, and which one your tenant is on Both listeners support two models, and the choice is not free - it is decided by whether your tenant uses database RBAC. | Model | What the client sends | Use it when | | --- | --- | --- | | shared | One listener-wide password. Every session arrives with the same identity. | Your tenant has no database RBAC grants at all. A shared credential cannot enforce per-user grants, so the engine will not let you combine the two. | | per-user | Each database user authenticates with its own password verifier, and its own grants apply. | Your tenant has any database RBAC grant. This is the only model that can serve it. | ### Two refusals you should expect, and what they mean These are designed behaviour, not faults. The engine refuses to bind a listener whose rules it could not enforce, rather than binding one that quietly serves more than it should. 1. A shared credential against a tenant that has RBAC: ``` DENIED: this tenant has database RBAC configured on N object(s) (...). ANY grant counts ``` Remedy: switch the listener to per-user login. Any grant counts - one grant on one object is enough to make a shared credential inadmissible, and so does any row-level security policy; a row policy is not the lighter case. Note what that buys and what it costs: the session becomes enforced, and [read-only](https://originchaindb.com/docs/connect-wire-protocols#per-user-read-only). 2. Per-user login against a tenant with no usable database user: ``` DENIED: per-user login is configured but this tenant has NO enabled database user with a password verifier, so no session could ever be admitted. Refusing to bind. ``` Remedy: create at least one enabled database user that has a password verifier. ``` -- Give the tenant at least one ENABLED database user WITH a password -- verifier, and the per-user listener has something it can admit. CREATE USER app_service WITH PASSWORD 'generate-a-real-secret-here'; GRANT SELECT, INSERT, UPDATE ON invoices TO app_service; ``` per-user login is read-only - on both listeners A per-user session refuses every write, on the MySQL listener and on the TDS listener alike. Not only the writes a grant disallows - the listener marks the session read-only the moment it admits an authenticated identity, before any grant is consulted, because there is no write-parity bridge on either adapter. `INSERT`, `UPDATE`, `DELETE` and DDL all come back refused, and no configuration turns this on. Read that together with the rule above - any grant at all forces per-user - and the consequence is the one to plan around: a tenant with a single RBAC grant or a single row-level security policy cannot write over either wire protocol at all. It must be on per-user login, and per-user login does not write. There is no combination of settings that yields an authorizing, writing session on these two doors today. Every write shown on this page therefore assumes a shared-credential listener - the one case the shared model is admitted in, which is also the only case in which these doors write. If your tenant has any authorization to enforce, keep these listeners for reads and write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client), where writes under an enforced identity are served. Put the two together and there is a real state that neither model can serve: a tenant that has database RBAC or a row-level security policy but has no database user with a password verifier. Shared login is refused because that authorization is present; per-user login is refused because there is no account to admit. This is not an exotic corner - it is the ordinary state of a tenant whose access has only ever been through API tokens. Adding one database user with a password is the whole fix. ## MySQL wire protocol Point any client that speaks the MySQL wire protocol at `$OC_SQL_HOST:3306`, authenticate with your tenant's shared or per-user credential, and issue SQL. TLS is not optional here The MySQL adapter fails open on TLS - but only on the shared credential model. Configured with a shared credential and no certificate it still binds, with TLS not required - and then every query and every result row crosses the network in cleartext, including columns your masking policy rewrites and rows your row-level security filters. A per-user listener is the opposite: it refuses to start at all without a certificate, because that model reads the cleartext password off the session and will not do so unencrypted. So the fail-open hazard below is a shared-credential hazard; per-user login has TLS settled for you. Those protections are applied inside the engine; they do nothing about an observer on the wire. So a MySQL listener must never be enabled without a certificate, and your client should demand TLS rather than merely accept it - `--ssl-mode=REQUIRED` for the CLI, or `ssl` with `rejectUnauthorized: true` for `mysql2`. The TDS listener behaves the opposite way and [fails closed](https://originchaindb.com/docs/connect-wire-protocols#sqlserver). ### Worked example - Sequelize 6.37 over mysql2 3.24 ``` // Verified against the engine with Sequelize 6.37 on mysql2 3.24. import { Sequelize, DataTypes } from "sequelize"; const sequelize = new Sequelize( process.env.OC_DATABASE, process.env.OC_DB_USER, process.env.OC_DB_PASSWORD, { host: process.env.OC_SQL_HOST, port: 3306, dialect: "mysql", dialectModule: require("mysql2"), dialectOptions: { // Required in practice. See "TLS is not optional here" below. ssl: { minVersion: "TLSv1.2", rejectUnauthorized: true }, }, // The engine cannot reflect a schema yet, so DEFINE every model // explicitly and never call sequelize.sync() to discover one. define: { timestamps: false, freezeTableName: true }, }, ); const Invoice = sequelize.define("invoices", { id: { type: DataTypes.INTEGER, primaryKey: true }, customer_ref: DataTypes.STRING, amount_cents: DataTypes.INTEGER, status: DataTypes.STRING, }); await sequelize.authenticate(); // Single-table reads and writes work, transactions included. const tx = await sequelize.transaction(); try { await Invoice.create( { id: 4101, customer_ref: "ACME-77", amount_cents: 129900, status: "open" }, { transaction: tx }, ); await Invoice.update({ status: "settled" }, { where: { id: 4101 }, transaction: tx }); await tx.commit(); } catch (err) { await tx.rollback(); // Rollback is honoured; the row does not appear. throw err; } const open = await Invoice.findAll({ where: { status: "open" }, order: [["amount_cents", "DESC"]], limit: 50, }); ``` ### What this does not do Measured against the real engine, on a listener using the shared credential model: a full ORM connects, authenticates, and runs a complete single-table workload - reads, writes, and transactions that both commit and roll back. On a per-user listener the writes in that sentence do not happen at all - see [below](https://originchaindb.com/docs/connect-wire-protocols#per-user-read-only). Two further things it cannot do on either model: - Schema reflection. The ORM cannot read table shapes back from the listener, so anything that discovers a schema - `describeTable`, `sync`, model autoloading, and the migration tooling built on them - does not work. Declare your models by hand. - Associations. Related models do not load. Eager loading with `include` and the lazy accessors that go with it do not resolve. do not call sync() A failed `sync()` is not a no-op. DDL is never buffered on this listener - `CREATE`, `DROP`, `ALTER`, `TRUNCATE` and `RENAME` go straight to the store instead of through the transaction buffer, and an implicit block that was open is force-committed first so it cannot be discarded behind them. That is the mechanism, not an implicit commit arriving afterwards: the `CREATE TABLE` statements a sync has already issued stay applied when the reflection step it runs next fails, and a `ROLLBACK` will not take them back. You are then left with a half-built schema and an exception that does not say so. Create tables deliberately - through the [schema API](https://originchaindb.com/docs/schemas) or explicit DDL - and keep the ORM to reads and writes. The practical consequence is worth stating plainly: this suits a single-table application, not a modelled domain. If your application's value is in its object graph, the MySQL listener is not the right door - use the [SQL endpoint](https://originchaindb.com/docs/sql) or the PostgreSQL wire, where joins are served. Fixes for both gaps are in flight. They are not shipped, and you should not plan against them. ``` // NOT served today. Both of these fail against the MySQL listener. // 1. Schema reflection - there is nothing to read a table shape back from. await sequelize.getQueryInterface().describeTable("invoices"); // fails await sequelize.sync(); // fails // 2. Associations - a modelled domain does not load. Invoice.hasMany(LineItem, { foreignKey: "invoice_id" }); await Invoice.findAll({ include: [LineItem] }); // fails // Do this instead: declare every model by hand (as above) and issue the // second query yourself, joining in your own application code. const invoices = await Invoice.findAll({ where: { status: "open" } }); const items = await LineItem.findAll({ where: { invoice_id: invoices.map((i) => i.id) }, }); ``` ## TDS - the SQL Server wire protocol Point a TDS client at `$OC_SQL_HOST:1433`. The same two credential models apply. TLS is mandatory, and it fails closed A certificate and key are required. Without loadable PEMs the port is not bound at all - there is no cleartext fallback and no degraded mode to ship by accident. If the port does not answer, an unloadable certificate is the first thing to check. Set `encrypt=true` and leave `trustServerCertificate` at `false` so your client actually validates the chain. ### Worked example - mssql-jdbc 13.4.0 and 12.10.1 ``` // Verified with mssql-jdbc 13.4.0 and 12.10.1 - 34 of 34 checks passed. String url = "jdbc:sqlserver://" + System.getenv("OC_SQL_HOST") + ":1433" + ";databaseName=" + System.getenv("OC_DATABASE") + ";encrypt=true" // mandatory - the listener is TLS-only + ";trustServerCertificate=false" + ";hostNameInCertificate=" + System.getenv("OC_SQL_HOST") + ";loginTimeout=30"; try (Connection conn = DriverManager.getConnection( url, System.getenv("OC_DB_USER"), System.getenv("OC_DB_PASSWORD"))) { // Prepared statements, including server-side handles and reuse. try (PreparedStatement ps = conn.prepareStatement( "SELECT id, customer_ref, amount_cents FROM invoices" + " WHERE status = ? ORDER BY amount_cents DESC")) { ps.setString(1, "open"); try (ResultSet rs = ps.executeQuery()) { while (rs.next()) { System.out.println(rs.getInt("id") + " " + rs.getString("customer_ref")); } } } // Batch execution. conn.setAutoCommit(false); try (PreparedStatement ins = conn.prepareStatement( "INSERT INTO invoices (id, customer_ref, amount_cents, status)" + " VALUES (?, ?, ?, ?)")) { for (Object[] row : rows) { ins.setInt(1, (Integer) row[0]); ins.setString(2, (String) row[1]); ins.setInt(3, (Integer) row[2]); ins.setString(4, (String) row[3]); ins.addBatch(); } ins.executeBatch(); } conn.commit(); } ``` ### Worked example - microsoft/go-mssqldb 1.11.0 ``` // Verified with microsoft/go-mssqldb 1.11.0 - 21 of 21 checks passed. package main import ( "context" "database/sql" "fmt" "net/url" "os" _ "github.com/microsoft/go-mssqldb" ) func main() { dsn := (&url.URL{ Scheme: "sqlserver", User: url.UserPassword(os.Getenv("OC_DB_USER"), os.Getenv("OC_DB_PASSWORD")), Host: os.Getenv("OC_SQL_HOST") + ":1433", RawQuery: url.Values{ "database": {os.Getenv("OC_DATABASE")}, // Mandatory - the listener is TLS-only and validates. "encrypt": {"true"}, "TrustServerCertificate": {"false"}, }.Encode(), }).String() db, err := sql.Open("sqlserver", dsn) if err != nil { panic(err) } defer db.Close() rows, err := db.QueryContext(context.Background(), "SELECT id, customer_ref, amount_cents FROM invoices"+ " WHERE status = @p1 ORDER BY amount_cents DESC", sql.Named("p1", "open")) if err != nil { panic(err) } defer rows.Close() for rows.Next() { var id, amount int var ref string if err := rows.Scan(&id, &ref, &amount); err != nil { panic(err) } fmt.Println(id, ref, amount) } } ``` ### How far this has been proven The driver-level defects that previously blocked both commercial drivers are fixed, and interop is proven end to end against the real engine - the framed TLS handshake, prepared-statement handles and batch execution included. The measured results were 34 of 34 for mssql-jdbc (13.4.0 and 12.10.1) and 21 of 21 for microsoft/go-mssqldb 1.11.0. One qualification we would rather state than have you discover: that interop battery runs in no automated job. It was executed deliberately and it passed, so the accurate word is verified, not continuously verified. A regression introduced by a future engine build would not be caught by a scheduled run today, so re-run the battery against the build you intend to deploy rather than trusting the figures above to still hold. And a second, narrower one, because it lands exactly where the [credential models](https://originchaindb.com/docs/connect-wire-protocols#credentials) do: those driver runs were made against a listener on the shared credential model, on a tenant with no database RBAC - the one case the shared model is admitted in. Per-user login has not yet been driven by either commercial driver; it is exercised only by our own test client, and we are not going to tell you it is proven when the drivers that matter to you have not been pointed at it. It also does not behave the same way, and in one respect the difference is decisive rather than incidental: a per-user session is [read-only](https://originchaindb.com/docs/connect-wire-protocols#per-user-read-only), on both the MySQL and the TDS listener. So the driver figures above were measured on a model that writes, against a model that does not. If your tenant has RBAC or a row-level security policy - and so must be on per-user - do not read those numbers across; pilot the login leg first, and plan your writes onto another door. ## Applications written for Oracle Database dialect, not transport This distinction decides whether you can adopt us, so it is worth being blunt about. We support a slice of the Oracle SQL dialect, and we serve it over the PostgreSQL wire. We do not implement Oracle Database's TNS network protocol, there is no listener on `1521`, and we do not intend to build one. That is a decision, not a backlog item - the compatibility matrix answers no, not later. The consequence: an Oracle client library cannot connect to us at all. An Oracle-shaped application connects through a PostgreSQL driver, and what has to port is the SQL, not the transport. If your application can be repointed at a PostgreSQL driver, a large slice of its SQL already works unchanged. If it is welded to an Oracle client library, it cannot connect, and no amount of dialect coverage changes that. ``` # An application written for Oracle Database connects with a PostgreSQL # driver, on the PostgreSQL wire port. There is no TNS listener to point # an Oracle client library at. psql "host=$OC_SQL_HOST port=5432 dbname=$OC_DATABASE user=$OC_DB_USER sslmode=require" ``` ### Served, and equivalent within scope | Construct | Scope | | --- | --- | | DUAL | The one-row dual table. | | NVL | Two-argument form; the operands must share one type. An empty string is not NULL here. | | NVL2 | Three-argument form; the branches must share one type. An empty string is not NULL here. | | Sequences | CREATE SEQUENCE and the sequence object itself. | | NEXTVAL pseudocolumn | seq.NEXTVAL as a projection - the Oracle spelling. | | MINUS | A pure spelling alias for EXCEPT: the same operator, whole-word rewritten. | ### Served, but narrower than Oracle Database Each of these ships and is useful, and each is narrower than the construct you know. Where a shape falls outside the served scope it is refused by name rather than executed with the clause quietly dropped - a dropped clause would return a confident, wrong answer. | Construct | Where it narrows | | --- | --- | | CONNECT BY | START WITH ... CONNECT BY [PRIOR] col = col over a single relation, rewritten to the equivalent WITH RECURSIVE. The hierarchy walk is served, and so are the pseudocolumns LEVEL, CONNECT_BY_ROOT col and SYS_CONNECT_BY_PATH(col, sep) - each carried down the recursion, so WHERE LEVEL <= 3 post-filters the walked hierarchy exactly as Oracle Database does. Refused by name: CONNECT_BY_ISLEAF, ORDER SIBLINGS BY, a PRIOR outside the CONNECT BY condition, more than one relationship, anything but a single PRIOR col = col equality, joins or multiple tables, a subquery in the FROM, and GROUP BY / HAVING / DISTINCT. A pseudocolumn is also refused as a sort key in every spelling - sort on a plain column instead. Nothing is executed with the clause dropped: that would return the unwalked base table with a 200. | | ROWNUM | The row-cap idioms port, and Oracle Database's counter semantics are preserved rather than approximated. WHERE ROWNUM <= n becomes a row cap; a projected ROWNUM rn numbers the returned rows from 1; the nested pagination idiom (... WHERE ROWNUM <= 20) WHERE rn > 10 returns the window Oracle returns; and WHERE ROWNUM > 1 matches nothing, as in Oracle, rather than being lowered to an offset. Refused by name, because ROWNUM is assigned before the sort and before grouping: a cap in the same block as ORDER BY, or alongside DISTINCT / GROUP BY / HAVING / an aggregate / a window function; a cap on a level that already carries LIMIT / OFFSET / FETCH; and ROWNUM under an OR, compared to a column, or used in GROUP BY / HAVING / ORDER BY / a join condition. For top-n write ORDER BY x FETCH FIRST n ROWS ONLY. One inconsistency to know about: the capability payload still folds ROWNUM into its rowid entry and reports it as roadmap, because that entry covers both pseudocolumns and has not been split yet. The behaviour above is what the engine does. | | DECODE | NULL equals NULL, and there is no implicit conversion of the search terms. | | INSTR | Two-argument form only; the position form is refused. | | SUBSTR | PostgreSQL semantics, positive offsets only. | | MERGE | Ships, and is narrower than Oracle Database's full MERGE grammar. | | (+) outer join | One provably equivalent shape: a single SELECT over exactly two comma-joined tables, all marks on one side, each marked predicate a plain col = col. Every other shape is refused rather than answered as an inner join, which would silently lose the unmatched rows the query was written to keep. | | Recursive queries | WITH RECURSIVE, within a documented narrower scope. | | SYSDATE | Ships, narrower than Oracle Database. | | TO_DATE | Ships, narrower than Oracle Database. | ### Refused today, with the portable rewrite | Construct | What to write instead | | --- | --- | | ROWID | Refused. Row identity here is the declared PRIMARY KEY, which every registered table has. A ROWID in Oracle Database is a physical address and a primary key is not one, so code that persisted ROWIDs was relying on something a primary key does not provide - that logic needs revisiting rather than translating. The row-cap uses of ROWNUM are a different construct and do port; see the preview table above. | | PIVOT / UNPIVOT | Refused. Build the cross-tab with conditional aggregates - SUM(CASE WHEN ... THEN ... END) with a GROUP BY - which is served over a single table and over a join. | | Flashback query | Refused: AS OF TIMESTAMP, AS OF SCN, VERSIONS BETWEEN. A historical read here is an operator-driven point-in-time restore at the backup surface, not a clause inside a SELECT. | | PL/SQL | Refused. Stored program units are not accepted. | | TNS network protocol | Not implemented, and not planned - the compatibility matrix answers no, not later. See the callout above. | ``` -- Served, and equivalent within the documented scope: SELECT NVL(discount_pct, 0), NVL2(shipped_at, 'shipped', 'pending') FROM orders; SELECT seq_invoice.NEXTVAL FROM DUAL; SELECT sku FROM catalogue MINUS SELECT sku FROM discontinued; -- Served, but narrower than Oracle Database (see the scope column): SELECT LEVEL, id, manager_id FROM staff START WITH manager_id IS NULL CONNECT BY PRIOR id = manager_id; SELECT DECODE(status, 'O', 'open', 'C', 'closed', 'other') FROM invoices; MERGE INTO invoices t USING staging s ON (t.id = s.id) WHEN MATCHED THEN UPDATE SET t.status = s.status; -- ROWNUM row caps port, with Oracle's counter semantics kept: SELECT id FROM invoices WHERE ROWNUM <= 10; -- becomes a row cap SELECT id FROM invoices WHERE ROWNUM > 1; -- returns NOTHING, as in Oracle SELECT * FROM ( -- the pagination idiom, verbatim SELECT a.*, ROWNUM rn FROM ( SELECT id, ref FROM invoices ORDER BY id ) a WHERE ROWNUM <= 20 ) WHERE rn > 10; -- REFUSED by name, because ROWNUM is assigned BEFORE the sort: this is -- ten arbitrary rows THEN sorted in Oracle, not the top ten by date. -- SELECT * FROM invoices WHERE ROWNUM <= 10 ORDER BY created_at; SELECT * FROM invoices ORDER BY created_at FETCH FIRST 10 ROWS ONLY; -- NOT served. ROWID has no equivalent; row identity is the PRIMARY KEY. -- SELECT ROWID, id FROM invoices; -- refused ``` ## A note on the names used on this page Every third-party name here is used to describe protocol or dialect compatibility, and for no other purpose. MySQL and Oracle are registered trademarks of Oracle Corporation and/or its affiliates. Microsoft and SQL Server are trademarks of the Microsoft group of companies. PostgreSQL is a registered trademark of the PostgreSQL Community Association of Canada. Sequelize, mysql2, mssql-jdbc and go-mssqldb are the property of their respective owners. OriginChainDB is not affiliated with, endorsed by, or sponsored by any of them. OriginChainDB is not any of these products - it is an independent database that speaks their wire protocols and accepts parts of their SQL. --- # Dashboard walkthroughs - OriginChainDB AI-native database Canonical source: https://originchaindb.com/docs/dashboard Sitemap last modified: 2026-09-23T08:33:23.000Z use the dashboard # Use the dashboard. Use the console to explore data, write queries, inspect relationships, and manage an instance. The current Workbench includes SQL, Vector, Full-text, and Graph modes. Screenshots in these guides use local sample data; live engine behavior is documented separately in the API and SDK references. [Create an instance → Choose a free or dedicated configuration, review the resources, and distinguish a local preview from live provisioning.](https://originchaindb.com/docs/dashboard/create-instance)[Navigate your database → Find your instance, understand the workspace sidebar, and open connection and operations settings.](https://originchaindb.com/docs/dashboard/navigate)[Create a schema → Define native fields in the table builder, inspect connections, and configure indexes separately.](https://originchaindb.com/docs/dashboard/schema)[Insert data → Validate a CSV or JSON batch, add individual records, and verify the resulting data.](https://originchaindb.com/docs/dashboard/insert)[Run queries → Choose SQL, vector, full-text, or graph inputs and inspect the returned records.](https://originchaindb.com/docs/dashboard/queries)[Workbench → Use the query library, editor tabs, result formats, saved drafts, and local history.](https://originchaindb.com/docs/dashboard/workbench)[Explain and plans → Open Explain from the Workbench and use Query stats or Slow queries to investigate costly work.](https://originchaindb.com/docs/dashboard/explain) Looking for the HTTP API? See the [API reference](https://originchaindb.com/docs/api). Looking for the SDK? See [SDKs](https://originchaindb.com/docs/sdk). Backups, point-in-time recovery and failover live in the [ops runbook](https://originchaindb.com/docs/ops). --- # Create an instance - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/create-instance Sitemap last modified: 2026-09-23T08:33:23.000Z Getting started # Create an instance Configure resources, review the deployment, and distinguish a local preview from a live instance. Configure an instanceOriginChainDB console [[Image: Current instance creation page showing the configuration step and local preview context before review.]](https://originchaindb.com/_astro/launch.BWiyIGxh.png) Configure an instance in the current console; the captured workflow saves a local preview.Current console · Dark theme · Local preview ## 1. Start from your instance inventory [Sign in to the console](https://app.originchain.ai/signin), open Instances, then choose New instance. The current dialog has two steps: Configure and Review. A page labeled Local preview creates a browser entry, not infrastructure. It does not issue production credentials or start billing. Use a live quote and provisioning response to confirm a real deployment. ## 2. Choose free or dedicated resources Free provides a simple starting configuration. Dedicated opens compute, storage, availability, and feature controls. The current free preview displays shared compute and 500 MB of local preview storage; confirm the live plan's quota and eligibility in the connected console. Supply an instance name and environment, then choose a cloud provider and region. The current preview accepts 3?40 lowercase letters, numbers, or hyphens, starting with a letter, and checks for duplicate names in the organization. Live deployment validation is authoritative. ## 3. Review a free configuration With Free selected, review the supplied compute, storage, and feature summary. You do not configure dedicated resources in this mode. A local preview can demonstrate the console without a billable database; a live free database requires a successful provisioning response. When using a managed free instance, check its current storage quota, idle behavior, and per-account eligibility. These service limits are separate from the browser preview's sample configuration. ## 4. Configure a dedicated instance | Setting | What to review | | --- | --- | | Compute | The selected vCPU and memory per node. | | Storage | Capacity, storage mode, and any cache allocation. | | Availability | Node count and the topology supported by the selected configuration. | | Recovery | The requested retention window and any recovery feature prerequisites. | | Networking | Public or private connectivity requirements. | | Capabilities | SQL, vector, graph, full-text, and optional features supported by the instance. | | Release channel | The engine release offered for the target deployment. | Options marked Preview, including a preview topology, must not be treated as a production availability guarantee. Saved feature preferences do not activate engine capabilities. A live quote is required to confirm price, region availability, and valid combinations. Read [deployment options](https://originchaindb.com/docs/deploy) and [replication limits](https://originchaindb.com/docs/replication/multi-region) before choosing the topology for a production workload. ## 5. Review before creating Choose Review instance and check the name, region, resources, features, and network request. Use Edit configuration to change them. In local preview, Create preview instance saves only the browser entry. For a live launch, review the quote and any required approval or payment step before submitting. Follow the actual provisioning status; an inventory row or a draft is not proof that the database is running. If credentials are shown once, store them securely. Rotating a credential can invalidate the previous value and affect connected clients. ## Next: connect and query Open [the instance overview](https://originchaindb.com/docs/dashboard/navigate), use Connect for your application or client, then follow the [quickstart](https://originchaindb.com/docs/quickstart) against a confirmed live endpoint. --- # Inspect query plans and diagnostics - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/explain Sitemap last modified: 2026-09-23T08:33:23.000Z Console guide # Inspect query plans and diagnostics Find costly query patterns, inspect their plans, and compare changes from the Workbench. Inspect a query planOriginChainDB console [[Image: Current Workbench showing the Query plan result tab for a sample SQL query, with the full application shell.]](https://originchaindb.com/_astro/explain.BOVhkJkx.png) The integrated Query plan view in the current Workbench, using an illustrative local plan.Current console · Dark theme · Local preview ## 1. Open Explain from the Workbench In Workbench ? Editor, select SQL and prepare a query. Open the arrow beside Run, then choose Explain plan or Explain and run. The plan appears in the results area's Query plan tab. Explain and run executes the query. Use plan-only mode when you only need its structure. Keep exploratory queries bounded while iterating; a LIMIT at the top of a plan does not necessarily prevent work below it. The local preview displays schematic operators and sample data. Its browser durations are not live engine measurements, and per-operator timings appear only when collected. ## 2. Read from inputs to output The operator at the top produces the final result. Follow its inputs down to see how rows are read, filtered, and joined, then read upward to understand the data flow. Select an operator to inspect its details. | Operator | What to inspect | | --- | --- | | Scan | The table and rows read; on a live plan, inspect any segment-scan or pruning details it reports. | | Filter | The predicate and how many input rows survive it. | | Join | The join keys and work performed by each input. | | Limit | The output cap, while checking how much work happened before it. | On a live analyzed plan, self time excludes child operators. Compare actual rows, estimated rows, and self time only when the response provides them; an absent value is not zero. ## 3. Keep plans and execution separate Explain plan does not execute the query. It can help you inspect the intended operators, but it cannot tell you the actual elapsed time or rows processed. Explain and run executes the query and may return measured values from the connected engine. ## 4. Find expensive patterns in Query stats Review query statisticsOriginChainDB console [[Image: Current Workbench Query stats page showing browser-local sample executions and the instance navigation.]](https://originchaindb.com/_astro/stats.CrFgSshc.png) Query stats in the current Workbench preview, with sample activity labeled.Current console · Dark theme · Local preview Open the Workbench's Query stats tab. SQL literals are replaced with placeholders so similar executions can be grouped. Start with total time, then compare calls, latency percentiles, and average returned rows. Select a pattern to inspect it. Open in editor and Explain in editor prepare a new draft; neither runs it. Replace any placeholder values and review the query before execution. In local preview, the stats use recorded browser activity and optional sample activity. They are not a complete audit log or live telemetry. The classic engine registry is in memory, holds up to 200 fingerprints by total time, and resets on engine restart. ## 5. Inspect an individual slow execution Slow queries shows individual executions. Set the duration threshold, inspect a row's query and result details, and reopen it in the editor when you need to reproduce the work. Opening an entry never executes it automatically. Raw query text can contain sensitive literal values. Live access and masking policies still apply. The engine's recent-execution window is an in-memory ring of up to 200 entries that resets on restart; local preview history is a separate browser record. ## 6. Change one thing and compare - If a scan reads many rows that a filter discards, inspect the predicate and available secondary indexes. - If a join dominates, check its keys and whether each input can be narrowed earlier. - If you return more rows or columns than the application uses, reduce the projection or result limit. - If tail latency is high, reproduce a slow execution with its actual parameter values. Add or change indexes through [Data Explorer](https://originchaindb.com/docs/dashboard/schema) or the [schema API](https://originchaindb.com/docs/schemas/reference), then compare the same query again. An index is useful only if the resulting plan uses it for the workload. ## SQL example and plan shape These examples illustrate a join and its operator structure. The actual plan depends on the query, schema, and engine version. ``` SELECT o.id, c.name, c.country, o.amount_cents FROM shop.orders o JOIN shop.customers c ON o.customer = c.id WHERE o.status = 'paid' LIMIT 10 ``` ``` limit(10) hash_join(on customer = id) filter(Eq { path: "status", value: String("paid") }) scan(shop.orders) scan(shop.customers) ``` Continue with the [SQL reference](https://originchaindb.com/docs/sql) for supported query syntax and EXPLAIN behavior. --- # Database users and access (IAM) - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/iam Sitemap last modified: 2026-09-23T08:33:23.000Z # Database users and access Find identity settings and prepare network and row-policy drafts for an instance. The screenshots show the current local console. Drafts are stored in this browser; they do not create credentials, change network access, or enforce database permissions. ## 1. Open Access Control Select an instance, then choose Access Control in its sidebar. Roles & grants, Database identities, Policies, Column masking, Effective access, and Audit are grouped here. Organization membership is managed separately under Team. ## 2. Check database identities Open Database identities. The current preview reports that the live identity list and credential operations require integration. It does not generate database users, passwords, or API keys. Check identity availabilityOriginChainDB console [[Image: Current Access Control Database identities page showing that live identity and credential operations require integration.]](https://originchaindb.com/_astro/identities.jPZw9Jzp.png) Access Control shows the identity integration status for the selected instance.Current console · Dark theme · Local preview A console member and a database user are separate identities. Adding someone to Team does not establish their right to query records. ## Credentials for a connected instance Use Connect on an instance to find application and database-client setup. Replace template values with credentials from an authenticated service. A sample endpoint or copied template is not a working credential. Credential kinds in the service API These credential lifecycles describe the service API, not actions available in the local preview. | Credential | Scope | Lifecycle | | --- | --- | --- | | Shared database bearer `oc_live_…` | One per database, shared by everyone and every deployed client. | Rotate replaces it. Every client still using the old value stops working. | | Personal access token `oc_u_…` | Yours alone, one or many per database. Carries read/write or read-only. | Rotate or revoke affects only that token — nobody else's breaks. | | Database-user API key `issued per database user` | Belongs to a database user, so queries run as that user with its roles and grants. | “+ key” issues another; deleting the user removes them all. | A personal token can be rotated or revoked independently. Rotating a shared database bearer invalidates its previous value for every client using it. A database-user key acts with that user's roles and grants; a newly issued secret is shown only once by the issuing service. ## 3. Prepare network access From the instance overview, choose Network. Add an address rule using an IPv4 or IPv6 address or CIDR range. A single address is normalized to `/32` or `/128`. Choose the empty-allowlist behavior explicitly. Removing the last rule does not change that choice. Private endpoint and peering options depend on the deployment plan and provider capabilities. Save or export the draft for review; no cloud firewall changes are made by this preview. ## 4. Prepare row policies Open Access Control → Policies to inspect existing definitions or prepare a policy draft. Saving a definition does not enable row filtering. Column masking also requires a connected engine before enforcement can be inspected. Prepare a row policyOriginChainDB console [[Image: Current Access Control Row policies page with a clear notice that definitions are drafts and do not enable row filtering.]](https://originchaindb.com/_astro/policies.BRDphGJw.png) Row-policy definitions are clearly marked as drafts in the current console.Current console · Dark theme · Local preview Engine policy semantics The documented read-policy model attaches a `USING` predicate to a table and role. Its predicate surface includes `=`, `!=`, `<`, `<=`, `>`, `>=`, `IN (...)`, `IS [NOT] NULL`, `AND`, `OR`, `NOT`, and `CURRENT_USER`. For the documented database-user read path, a policy sees real values before column masking; rows failing the predicate are filtered out. A database bearer acting as the administrative superuser is not a test of restricted user access. With no policies on a table, the policy layer supplies no row filter. Verify identity propagation and enforcement separately for each query path before relying on a policy. ## Review access changes Access Control → Audit shows local role and policy draft edits. Filter the history by kind or text, then export the filtered entries. This history is not a durable security log or a verified engine audit chain. Connected engine audit reference The engine's authorization audit model is an append-only hash chain. Its documented actions include `user.create`, `user.delete`, `user.rotate_key`, `role.create`, `role.delete`, `grant.put`, `grant.revoke`, `rls.put`, and `rls.delete`. Each entry includes its target, change detail, and previous hash. Chain verification must recompute the links and identify any first broken sequence. Engine audit exports use NDJSON. Organization events and browser-local draft history are separate records; this preview does not fetch or verify the engine chain. ## Before applying a policy Connect an authenticated service and verify the deployed engine's capabilities. Local policy and role definitions establish no access boundary. | Control | Boundary to verify | | --- | --- | | Network access | The actual endpoint rules and empty-allowlist policy, independently of database credentials. | | Read-only token | Write requests are rejected by the service; a browser label alone cannot enforce this. | | Database identity | Queries propagate the intended database user rather than an administrative bearer identity. | | SELECT, policies and masking | Reads by that identity are checked for table access, row filtering, and masked columns on every supported query path. | | Insert, update, delete, references | The documented grant model records these privileges without promising a write barrier. Use a service-enforced read-only token or separate database when a write boundary is required. | Next: [prepare a role and inspect its proposed grants](https://originchaindb.com/docs/dashboard/rbac). --- # Add and import records - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/insert Sitemap last modified: 2026-09-23T08:33:23.000Z Console guide # Add and import records Validate CSV or JSON imports, add individual records, and verify the result in Data Explorer. Open Data ExplorerOriginChainDB console [[Image: Current console with Data Explorer selected, namespace and table controls, and sample document records.]](https://originchaindb.com/_astro/data.-VxrfpcF.png) Inspect table records and structure in the current Data Explorer preview.Current console · Dark theme · Local preview ## 1. Open a table and import records Import recordsOriginChainDB console [[Image: Import dialog over the current Data Explorer, with CSV and JSON options and local preview context.]](https://originchaindb.com/_astro/import.B3T423MX.png) Validate a CSV or JSON batch before importing it into the local preview.Current console · Dark theme · Local preview In Data Explorer ? Data, select a table and choose Import. The current dialog accepts CSV or JSON. Choose a file or paste its contents, then select Validate import. 1. Use field names as CSV headers, or provide a JSON array of records. 2. Review the validation result and the first records shown in the preview. 3. Fix any type, required-field, or key errors before importing. 4. Choose Import records to add the validated batch. The local preview accepts files up to 2 MB and allows at most 1,000 records per table, including existing rows. It validates the complete batch before changing the local dataset. ## 2. Resolve validation errors Check column names, native datatypes, required values, primary keys, and referenced records. Duplicate keys are rejected by the current local catalog. Do not assume that a live write endpoint has the same insert-versus-upsert behavior; check the [write API documentation](https://originchaindb.com/docs/insert). Empty optional CSV cells are omitted for native fields and become NULL for legacy fields. JSON integer tokens are parsed losslessly; exact integers export as strings in the local preview. Keep identifiers such as postcodes and zero-padded keys as text when their formatting matters. ## 3. Add a single record Use the Record button in the table's toolbar to enter one record. The form validates it against the current table definition. There is no need to use a SQL editor for a local preview record. The local Workbench's sample SQL is read-only. Use a connected, supported SQL endpoint or row API for production INSERT, UPDATE, and DELETE operations; the engine examples below are preserved for that workflow. ## 4. Check what was accepted A refused value should remain an error, rather than being treated as a successful write. Check for a wrong type, a misspelled field, an omitted required value, a duplicate key, or a violated reference. Engine CHECK constraints and endpoint-specific error codes are documented in [the error reference](https://originchaindb.com/docs/errors). ## 5. Verify the result Use the Data view to inspect records, filter fields, sort, or switch to JSON. Open in Workbench prepares a query for the selected table. Review the returned rows and their types; a saved browser record does not prove a production write succeeded. For a live import, check the response from the ingest endpoint and run a bounded read or count against the target table. Use `namespace.table` where the API requires a fully qualified name. ## Choose the API for larger imports The 2 MB file and 1,000-record limits above belong to the local console preview. For production batch writes, NDJSON, larger imports, vectors, or scheduled ingestion, use the [insert guide](https://originchaindb.com/docs/insert) and the live endpoint's documented limits. Do not assume a bulk import is a transaction. If a live upload fails after accepting rows, inspect its response before retrying. Row writes keyed by primary key may replace an existing row; choose an endpoint with the write semantics your application requires. Vector data uses the vector endpoint and its declared dimension and metric. ## Original CSV and SQL examples These examples target `shop.orders` from the [quickstart](https://originchaindb.com/docs/quickstart). They are live-engine examples, separate from the console's local sample parser. Show CSV, INSERT, response, and verification examples ``` id,customer,amount_cents,status,notes,placed_ms ord_3000,cus_481,1200,paid,,1753660800000 ord_A001,cus_9930,2111,paid,gift wrap,1753660801500 ord_H002,cus_2214,3022,paid,,1753660803000 ``` ``` INSERT INTO shop.orders (id, customer, amount_cents, status, notes, placed_ms) VALUES ('ord_7Q2R', 'cus_481', 15400, 'paid', 'expedited', 1753669000000) ``` ``` { "kind": "insert", "schema": "shop.orders", "rows": [ { "id": "ord_7Q2R", "customer": "cus_481", "amount_cents": 15400, "status": "paid", "notes": "expedited", "placed_ms": 1753669000000 } ] } ``` ``` SELECT COUNT(*) AS orders FROM shop.orders ``` Next, [query your data](https://originchaindb.com/docs/dashboard/queries) and inspect the returned records. --- # Navigate the console - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/navigate Sitemap last modified: 2026-09-23T08:33:23.000Z Console guide # Navigate the console Find data, queries, connections, and instance operations from one organized workspace. Find your instance toolsOriginChainDB console [[Image: Current OriginChainDB instance overview with its full workspace sidebar, header, and connection and configuration links.]](https://originchaindb.com/_astro/overview.BQSU2q0X.png) The current instance overview, showing local preview configuration and workspace navigation.Current console · Dark theme · Local preview ## 1. Pick an organization and instance The top bar identifies your organization and active instance. The instance workspace is where you query and manage that database; Back to instances returns to the organization inventory. Search and the theme control remain available across the console. Check the connection state before acting. Local preview means the page is showing sample data or configuration drafts, not confirmed production state. ## 2. Find the right workspace | Page | Use it to | | --- | --- | | Overview | Review the instance, deployment plan, connection entry point, and available telemetry. | | Workbench | Run queries, inspect results and plans, and review query stats or slow executions. | | Data Explorer | Browse rows, map connections, explore embeddings, and inspect indexes. | | Database Management | Inspect tables and database objects such as functions, triggers, and enumerated types. | | Access Control | Manage roles and policies for the selected instance. | | Configuration | Review compute, storage, availability, engine version, and feature settings. | | Platform | Find backup, replication, and migration workflows. | Controls that depend on an engine capability or a dedicated instance must be confirmed against that instance. A preview control is not a promise that a live change has been applied. ## 3. Browse your instances The organization-level Instances page is the inventory. Open an instance to enter its workspace, or choose New instance to prepare a deployment. Organization team, usage, billing, and settings pages apply across that organization. ## 4. Explore data and connect an application Data Explorer has four views: Data, Connections, Embeddings, and Indexes. Use namespace and table filters to narrow the scope. In Connections, switch between the schema map and record graph; Expand workspace gives the canvas more room without losing its controls. Open Connect from the instance overview or Workbench. Choose Application API, Database client, or AI & MCP, then follow the instructions for your client. Keep credentials private and confirm the live endpoint, enabled listener, and TLS settings before using a template. Network access, cold storage, configuration, and instance settings are linked from the overview. Dedicated-only settings are identified where you configure them. ## 5. Review changes before applying them A saved draft is different from an applied engine change. Read the resulting status, especially for network rules, feature switches, backup requests, and deployment settings. Deleting an instance or changing its credentials can affect connected applications; use the confirmation and impact details presented by the live console. ## Next: explore a query Open the [Workbench guide](https://originchaindb.com/docs/dashboard/workbench) for a first query, or [create a table](https://originchaindb.com/docs/dashboard/schema) before loading your own data. --- # Run queries in the Workbench - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/queries Sitemap last modified: 2026-09-23T08:33:23.000Z Console guide # Run queries in the Workbench Choose a query mode, run a bounded example, and inspect the returned records in the Workbench. Run a SQL queryOriginChainDB console [[Image: Current OriginChainDB console with Workbench selected in the sidebar, the SQL editor, and local sample query results.]](https://originchaindb.com/_astro/sql.33U5scMy.png) A SQL query and returned records in the current console preview.Current console · Dark theme · Local preview The current Workbench has SQL, Vector, Full-text, and Graph modes. Screenshots use local sample data. Live engine syntax and API support are documented in the references linked below. ## SQL: inspect structured rows 1. Open Workbench and choose SQL. 2. Start from Examples, or write a bounded SELECT against the displayed table. 3. Click Run and inspect Results. Switch the result format to JSON when you need its record structure. ``` SELECT id, title, category FROM documents LIMIT 5; ``` This query uses the local preview's sample documents. For a live query, use the table names and syntax required by your instance. The [SQL reference](https://originchaindb.com/docs/sql) covers joins, aggregates, limits, and supported statements. ## Vector: find similar records Choose Vector, enter the query embedding, set the number of results and optional category filter, then run it. The sample mode uses three dimensions and cosine similarity. A production query must match the stored vectors' dimension and configured distance metric. For a live embedding workflow, follow [the vector quickstart](https://originchaindb.com/docs/vector/quickstart). Use Data Explorer ? Embeddings to inspect a sample projection; a three-dimensional visualization does not replace the original embeddings or distance calculation. ## Full-text: find matching documents Choose Full-text, enter search terms, set the result count, and run. Ranked matches show the fields needed to inspect each hit. In the local preview, the search covers sample document titles and bodies. Run a full-text queryOriginChainDB console [[Image: Current console Workbench in Full-text mode with the sidebar, query input, and highlighted sample search matches visible.]](https://originchaindb.com/_astro/search.B60QkJ29.png) Full-text matches with highlighted terms from the local sample dataset.Current console · Dark theme · Local preview For live BM25, boolean, and phrase queries, use the [full-text reference](https://originchaindb.com/docs/fts) and verify that the relevant index exists. ## Graph: follow relationships Choose Graph and select a starting node, relationship, direction, and traversal depth. Run the traversal, then select a node or record to inspect its properties. The preview form offers one, two, or three hops over the sample graph. The live engine's supported Cypher syntax is a separate query surface; use [the Cypher guide](https://originchaindb.com/docs/schemas/cypher) and [graph API reference](https://originchaindb.com/docs/graph) for those requests. ## Natural language: use Ask Natural-language queries use the [Ask endpoint](https://originchaindb.com/docs/ask). The current Workbench preview does not expose an NL mode. Follow the Ask guide to scope schemas, inspect the compiled plan, and handle the returned rows in an application. ## Inspect results before reusing a query Choose a result format appropriate to the mode: table, JSON, ranked matches, or graph. Select a record for field details; export CSV when you need the returned records. The browser is not an unbounded data export tool, so keep queries limited and use an SDK for larger retrieval jobs. Save a useful draft from the toolbar, or reopen a previous input from History. Opening a history entry does not run it. See [Workbench in depth](https://originchaindb.com/docs/dashboard/workbench) for local persistence and [query plans](https://originchaindb.com/docs/dashboard/explain) for diagnostics. ## Engine examples for the shop dataset These original examples use `shop.orders` and `shop.customers` from the [quickstart](https://originchaindb.com/docs/quickstart). They document engine query surfaces, not the local sample parser. Create the example tables and use the corresponding supported endpoint before running them. SQL examples ``` SELECT id, customer, amount_cents, status FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 10 ``` ``` SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total_cents FROM shop.orders WHERE status = 'paid' GROUP BY customer ORDER BY total_cents DESC LIMIT 8 ``` ``` SELECT o.id, c.name, c.country, o.amount_cents FROM shop.orders o JOIN shop.customers c ON o.customer = c.id WHERE o.status = 'paid' LIMIT 10 ``` Cypher pattern ``` MATCH (o:orders {id: '01JTRX1H4Q9P0N2WMX0F5JZ001'})-[:customer]->(c) RETURN c.name, c.country LIMIT 10 ``` Ask question ``` which countries spent the most on paid orders ``` Full-text request examples ``` table:field one or more query terms shop.orders:notes rush delivery ^^^^^^^^^^^ ^^^^^ ^^^^^^^^^^^^^ table field terms ``` ``` shop.orders:notes rush delivery ``` --- # Roles and permissions (RBAC) - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/rbac Sitemap last modified: 2026-09-23T08:33:23.000Z # Roles and permissions Review organization membership and prepare scoped role drafts for a database. The current console keeps these changes in local preview storage. Team invitations do not send emails, and database role drafts do not grant or deny live access. ## 1. Review your team Return to the organization and open Team. Search members by name or email. Invite member accepts an email and a Developer, Admin, or Viewer role; Add demo invitation records an invitation locally. It does not create a live account or contact the recipient. Review your teamOriginChainDB console [[Image: Current organization Team page with member roles and the notice that preview changes do not send invitation emails.]](https://originchaindb.com/_astro/team.BG6YFwwv.png) The current organization Team page identifies its changes as local preview data.Current console · Dark theme · Local preview ## 2. Open Roles & grants Select an instance and open Access Control → Roles & grants. Search drafts or filter by their definition namespace. Existing text definitions remain intact; structured drafts can be inspected and edited. Open Roles and grantsOriginChainDB console [[Image: Current Access Control Roles and grants page showing local role drafts and the full instance sidebar.]](https://originchaindb.com/_astro/roles.XPi-ZDL7.png) Roles & grants lists access proposals for the selected instance, not live database grants.Current console · Dark theme · Local preview ## 3. Prepare a role draft 1. Choose New role draft and enter a name, definition namespace, and optional description. 2. Choose Add grant. Select its namespace, table, and query shape. 3. Select the proposed permissions: `select`, `insert`, `update`, `delete`, or `references`. 4. Choose Save role draft to keep the proposal locally, or Export draft to review its JSON outside the console. Prepare a role draftOriginChainDB console [[Image: New role draft dialog showing a proposed namespace, table, query shape, and permission grant over the current Access Control page.]](https://originchaindb.com/_astro/role-editor.CvEURcMg.png) A role proposal scopes each grant by namespace, table, query shape, and permission.Current console · Dark theme · Local preview All tables includes future tables in that namespace in the draft model. A grant must contain at least one permission. Duplicate namespace/table/query-shape scopes must be combined; removed tables or incompatible query shapes need correction. Structured drafts support up to 100 grants. ## 4. Review proposed effective access Open Effective access, select a structured role draft and query shape, then inspect the table matrix. It shows Proposed or Not proposed for the five permission types. This analysis combines direct grants only. It does not evaluate database users, inherited roles, row policies, column masks, or live engine permissions. A text definition cannot be interpreted by the structured analyzer, and Not proposed does not mean the database denies access. A database role belongs to the database identity layer. Organization Team membership does not provide database permissions. See [database identities](https://originchaindb.com/docs/dashboard/iam#database-users) and [the enforcement boundary](https://originchaindb.com/docs/dashboard/iam#enforced) before applying a proposal. ## Edit or remove a draft Return to Roles & grants and choose Edit draft. Removing a draft requires typing its name. These actions update local definitions and local Access audit history only; they do not alter a deployed role or its memberships. Organization role reference for a connected service The service role model below is distinct from the preview Team labels. Confirm the authenticated service's role catalog and permission checks before assigning production access. | Capability | Org owner | Org admin | Cluster operator | Read-only | Billing-only | | --- | --- | --- | --- | --- | --- | | Manage members & invites | Included | Included | Not included | Not included | Not included | | Create / delete clusters | Included | Included | Not included | Not included | Not included | | Operate scoped clusters | Included | Included | Included | Not included | Not included | | Edit billing & payment | Included | Included | Not included | Not included | Included | | View invoices | Included | Included | Included | Included | Included | | Read org audit log | Included | Included | Included | Included | Not included | | Manage API tokens | Included | Included | Not included | Not included | Not included | | Configure SSO · v2 | Included | Not included | Not included | Not included | Not included | | Read-only console access | Included | Included | Included | Included | Included | | Capability | Documented status | Service boundary | | --- | --- | --- | | Manage members & invites | Enforced by the service | Sending and revoking invitations, adding a member directly, resetting a member's password, changing a member's role, removing a member, and creating a custom role. All refuse with a permission error without it. | | Create / delete clusters | Enforced by the service | Deleting a database, and deleting a schema. Both refuse without it. | | Operate scoped clusters | Enforced by the service | Rotating a database's shared bearer token, and running schema migrations. Both refuse without it. | | Edit billing & payment | Descriptive | Descriptive. The billing pages are not gated on this capability today. | | View invoices | Descriptive | Descriptive. Invoice and usage views are not gated on this capability today. | | Read org audit log | Descriptive | Descriptive. The org audit log is not gated on this capability today. | | Manage API tokens | Descriptive | Descriptive. Personal access tokens are already per-member and scoped to the holder — see the IAM page. | | Configure SSO · v2 | Descriptive | Reserved for single sign-on, which is not shipped yet. Owner-only in the matrix. | | Read-only console access | Descriptive | The baseline every role holds; it is what lets a member open the console at all. | In this documented service model, custom organization roles compose the same nine capabilities; unresolved roles fall back to read-only console access. The owner safeguards refuse demoting or removing the last owner, and refuse removing yourself. A signed-in role alone does not establish a hard boundary for capabilities marked descriptive. Database grants use Select, Insert, Update, Delete, and References. The documented administrative role is `ll_admin`; system roles cannot be deleted. Read access, inherited roles, wildcard objects, and masked columns must be evaluated by the engine. The remaining write-grant letters must not be assumed to enforce writes; the [IAM reference](https://originchaindb.com/docs/dashboard/iam#enforced) explains this limit. ## Apply through an authenticated service Review exported scopes, resolve database identities, and verify enforcement for SQL, vector, full-text, graph, and direct record requests. A browser draft is not evidence that a user is restricted. For the documented service's offboarding flow, personal tokens are revoked by their holder; removing organization membership does not itself delete those tokens. Review shared credentials and database users separately. Keep another authorized owner available before changing administrative membership. Next: [prepare row policies and check the live enforcement boundary](https://originchaindb.com/docs/dashboard/iam#rls). --- # Create tables and configure indexes - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/schema Sitemap last modified: 2026-09-23T08:33:23.000Z Console guide # Create tables and configure indexes Define native fields, inspect relationships, and configure indexes from Data Explorer. Create a tableOriginChainDB console [[Image: New table dialog over the current Data Explorer, with the application header and sidebar visible behind it.]](https://originchaindb.com/_astro/schema.DKnE1r4c.png) The native field builder in the current console, using a versioned local datatype catalog.Current console · Dark theme · Local preview ## 1. Open Data Explorer Select your instance, open Data Explorer, and choose a namespace. Data lists tables and records; Connections shows the schema map or record graph. Choose New table to start a table definition. The current screenshots use a local catalog. Saving a preview table changes browser data; confirm the connection state before expecting a table to exist on a live instance. ## 2. Define native fields 1. Choose the namespace and table name, then add an optional description. 2. Add each field and choose its datatype from the loaded catalog. 3. Set Required and Primary key where the selected type permits them. 4. Complete type-specific settings, such as enum values or decimal parameters, when shown. 5. Review validation messages and choose Create table. The datatype picker follows the available catalog, rather than assuming every engine supports the same types. If it cannot load, keep the draft and use Retry; saving is disabled. In local preview, the catalog is explicitly labeled as a versioned local catalog. Primary-key fields are required. Optional native values may be omitted. The editor permits up to 40 fields in the local preview; that is a UI preview limit, not an engine capacity claim. ## 3. Configure indexes separately Inspect indexesOriginChainDB console [[Image: Current Data Explorer Indexes view with the application navigation and sample index catalog.]](https://originchaindb.com/_astro/indexes.Pfxpai-c.png) Inspect index fields and configuration in the current console preview.Current console · Dark theme · Local preview Open the Indexes view and scope it to your table. Select an index to inspect its fields, mappings, and configuration. Relational, vector, full-text, and graph indexes have different settings; selecting a table modality alone does not create a working vector or search index. Use the [vector schema guide](https://originchaindb.com/docs/schemas/vector) for dimensions, distance, and index family, or the [full-text guide](https://originchaindb.com/docs/schemas/fts) for searchable fields and analysis. A local configuration or requested job is not confirmation that a live engine index has been built. ## 4. Connect tables In Connections, switch between Schema map and Record graph. Namespace and table filters narrow the view. Expand workspace fills the available page while keeping filtering, inspection, and canvas controls nearby. Declare a relation before querying it through the graph API. Cross-namespace connections should use the fully qualified source and target table names required by the schema. ## 5. Inspect the table definition Open a table in Data to inspect its records, structure, references, and manifest. The preview's exported catalog manifest is JSON; it is not the same document as the engine's canonical TOML schema manifest. Use the API reference when preparing a live schema registration. ## 6. Edit a structure carefully Choose Edit structure from the table. Existing field names are fixed in the current editor, and populated field types remain unchanged. New fields and compatible parameter changes are validated against the stored preview records. ## 7. Remove dependencies before deletion Delete a table from its Structure view. Referencing relationships and indexes must be removed first, and the confirmation requires the table name. A live table deletion can remove data permanently; read the connected instance's confirmation before proceeding. ## Register an engine schema with TOML The original manifest examples below define `shop.orders`. They remain API examples: the current native field editor does not have a TOML editing tab. See the [schema reference](https://originchaindb.com/docs/schemas/reference) for fields, secondary indexes, relations, foreign-key actions, and CHECK constraints. Show the original TOML and registration examples ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" required = true [[columns]] name = "amount_cents" ty = "i64" required = true [[columns]] name = "status" ty = "str" required = true [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" [[indexes]] name = "by_status" columns = ["status"] [[check_constraints]] name = "amount_non_negative" expression = "amount_cents >= 0" ``` ``` [[relations]] name = "placed_by" from_col = "customer" target = { namespace = "shop", table = "customers", pk = "id" } bidirectional = true ``` Next, [add and validate records](https://originchaindb.com/docs/dashboard/insert), or follow the [schema tutorial](https://originchaindb.com/docs/schemas/tutorial) from an application. --- # Use the query Workbench - OriginChainDB Canonical source: https://originchaindb.com/docs/dashboard/workbench Sitemap last modified: 2026-09-23T08:33:23.000Z Console guide # Use the query Workbench Write queries, inspect results, and keep useful drafts together in one instance workspace. Run a SQL queryOriginChainDB console [[Image: Current OriginChainDB console with Workbench selected in the sidebar, the SQL editor, and local sample query results.]](https://originchaindb.com/_astro/sql.33U5scMy.png) The current Workbench, running a SQL example against local sample data.Current console · Dark theme · Local preview About this walkthrough The screenshots show the current console preview. Sample queries, saved drafts, and history run or persist in this browser; they are not evidence of a connected production database. Use the connection status shown in your console before running a live query. ## 1. Open an instance workspace Select an instance, then open Workbench in the sidebar. The page keeps the editor, results, and query library together. The tabs above it switch between Editor, Query stats, and Slow queries. ## 2. Choose what to explore The library selector opens Explorer, Saved, or History. Explorer shows fields and example data for the selected query mode. Use its search field to find a table or field. The editor's modality selector offers SQL, Vector, Full-text, and Graph. Each mode has its own input: SQL text, a query vector, search terms, or a graph traversal form. Start with Examples when trying a mode for the first time. [See the inputs each query mode expects →](https://originchaindb.com/docs/dashboard/queries) ## 3. Keep related queries in tabs Use New query to open a draft. The query-options menu can rename or duplicate it. Each tab retains its mode and input; an unsaved indicator marks changes. Closing an unsaved tab offers Save and close, Discard draft, or Cancel. The preview keeps up to 20 open query tabs per instance. Query drafts belong to the selected organization and instance in this browser. ## 4. Run and inspect results Click Run to execute the current input. The results area offers the formats relevant to that mode: a table, JSON, ranked matches, or a graph. Select a record to inspect its fields, or export the returned records as CSV. For SQL, the arrow beside Run opens Explain plan and Explain and run. The first inspects a plan without executing the query; the second executes it. A returned plan appears under Query plan beside Results. Plans in the local preview are schematic. Browser duration and illustrative operator values must not be read as live engine telemetry. ## 5. Save a query or revisit a run Reopen a queryOriginChainDB console [[Image: Current Workbench with History selected in its library and the application navigation visible.]](https://originchaindb.com/_astro/history.C4gBOYk_.png) Reopen a locally recorded input from the Workbench History library.Current console · Dark theme · Local preview Click Save to add a query to the Saved library. History records recent executions and lets you reopen their inputs. Opening an entry prepares a draft; it does not execute it. The preview stores up to 100 saved queries and 100 history entries locally. They do not sync between machines, and clearing browser storage removes them. A saved query is a bookmark, not a scheduled job. ## 6. Check the target before running The instance name in the workspace header identifies the current target. Use Connect for connection instructions, and check whether the page is showing local sample data or a live connection. Keep a `LIMIT` on exploratory SQL queries; the browser displays only the records returned to it. ## Continue in your application Move a working query into an [SDK](https://originchaindb.com/docs/sdk) or the [HTTP API](https://originchaindb.com/docs/api) when you need application credentials, bound inputs, or repeatable execution. Use [query plans and diagnostics](https://originchaindb.com/docs/dashboard/explain) to investigate costly work. --- # Deploy OriginChainDB - dedicated single-tenant configurations Canonical source: https://originchaindb.com/docs/deploy Sitemap last modified: 2026-09-23T08:33:23.000Z 06 · deploy # Deploy - what the managed configuration gives you. Choose a managed instance configuration and understand its region, storage, networking, and recovery options. Note for operators Most customers can skip ahead to [provisioning](https://originchaindb.com/docs/deploy#provisioning). The architecture and topology sections below are mainly informational - useful background on how your dedicated instance is laid out. ## Tenant architecture Free and Starter databases run on shared, scale-to-zero infrastructure; this page covers dedicated configurations. Each dedicated configuration is a managed instance for one tenant - primary plus an optional warm standby in the same region - fronted by a per-tenant managed wildcard cert and a single hashed bearer token. No shared disk, no shared memory. Account & billing (signup, console) is global; your data stays in the region you pick. Sharding, replicas, and high availability are engine-level concerns, not infrastructure tricks. A dedicated configuration runs one managed engine per tenant - nothing extra to schedule. ## Choosing a region + configuration | configuration | use | RAM | topology | | --- | --- | --- | --- | | Single-zone | First production workload, prototypes, dev environments | 8 GB | No SLA. Single instance. Card required at signup; billed from day 1. | | HA | Production SaaS with SLA | 16 GB | Warm standby for fault tolerance, 99.9% SLA. | | HA+ | High-concurrency customer-facing SaaS | 32 GB | Primary + 2 replicas for fault tolerance, 99.95% SLA. | | Enterprise | Custom - HIPAA BAA, GDPR DPA, dedicated capacity | by spec | Per-contract terms, 99.99% SLA target. | Pricing multiplier: the home region is the base price; every other region adds 1.15× on compute and storage to cover cross-region operational overhead. ## Selecting add-ons Base configurations ship with the core engine, single-row CAS, boolean FTS, and continuous-archive PITR. Everything else is an opt-in monthly add-on you can add or remove any time. See [/pricing](https://originchaindb.com/pricing) for line-item costs. SQL Pro Analytics workloads - JOIN, OUTER JOIN, GROUP BY, HAVING. Vector Search Embedding workloads - HNSW ANN with SIMD, filtered topk. Full-Text Pro Search workloads - BM25, phrase, Unicode, Snowball stemming. Graph Relationship workloads - neighbors, BFS, path, weighted Dijkstra. Transactions Multi-row OLTP - snapshot isolation with optimistic conflict detection. Intra-Segment PITR Tighter restore granularity via a continuous change archive. Preview - the archive ships on a configured interval; see ops → backups before you depend on a particular granularity. Multi-Writer Cluster Active-active cross-region replication. In development, not yet provisionable - contact sales. Add-ons attach to any configuration and bill on the next invoice. Toggle on or off from the console at any time; line items prorate to the day. ## Provisioning Click-to-running takes ~30 seconds. There are no manual steps; the console drives the whole flow. 1. Pick region + configuration Console → New instance. There is a home region at base price; other regions carry a 1.15× pricing multiplier on compute + storage. 2. Click create We provision a dedicated managed instance and, on high-availability configurations, a warm standby in the same region - both behind a TLS 1.3 listener with a managed wildcard cert under db.originchain.ai. 3. Copy your bearer Console mints a securely hashed bearer token at instance creation. One active token at a time. Rotation is a single click and emits an audit-log entry; the prior token is honored for 60s to cover rolling deploys. 4. Connect Endpoint is .db.originchain.ai. Median time from click to first 200 OK on /health is ~30 seconds. ## DNS & TLS Each tenant gets a DNS record auto-provisioned at `.db.originchain.ai` pointing at your instance. The managed wildcard cert under `*.db.originchain.ai` auto-renews; you never see the private key. When a standby is promoted, the same DNS record is re-pointed at the promoted instance with a 60-second TTL - propagation is typically sub-minute. ## Bearer rotation Rotate from the console at any time. The new token activates immediately and the prior token stays honored for 60 seconds so a rolling deploy can swap without a 401-storm. Every rotation lands an entry in the per-tenant audit log with actor, timestamp, and source IP. ## Replication topology High-availability configurations run a primary plus a standby in the same region, each on its own virtual machine; a single-zone configuration has no standby. A write is acknowledged once it is flushed to durable storage on the primary - a `2xx` means the data is on disk, not buffered in memory, and it survives the database process being killed. Recovering it afterwards has two caveats worth reading before you design around this, covering abrupt power loss and a full storage volume - see [ops → write durability](https://originchaindb.com/docs/ops#durability). Replication to the standby is asynchronous: the primary streams committed frames continuously but does not block on the standby. An abrupt loss of the primary can therefore cost the most recent acknowledged writes - the recovery point is not zero, and how much is at risk depends on how far behind the standby had fallen. Promotion of the standby is not unattended by default: an opt-in automatic-promotion path exists, and with it enabled the standby claims the primary role itself once the old primary's claim has expired, with an external restart hook bringing it back up as the writer. It refuses to promote a standby that is not fully caught up. With the option off, promotion is operator-driven. Either way the old primary's claim has to expire first, so it is not a failover to hold a time budget against (see [ops → failover](https://originchaindb.com/docs/ops#failover)). Split-brain is fenced by a single-primary claim, so only one instance ever accepts writes. Instances configured for synchronous acknowledgement go one step further on the row-write endpoints: the primary waits up to 500 ms for the standby before it replies. If the standby does not answer in that window the write still succeeds and the response carries `X-OC-Replication: degraded`. Treat that header, not the status code, as the signal that recent writes are at risk. The other write surfaces - SQL, Cypher, transaction commit, vector and full-text index writes - do not wait for the standby at all. Cross-region active-active replication is in development and is not yet available for provisioning - see [active-active](https://originchaindb.com/docs/replication/multi-region). Quorum commit is roadmap and not scheduled: no engine in service commits a write through consensus, every database runs the active-passive path, and a running database cannot today be switched into a multi-node consensus group. Today every byte stays in the single region you picked. ## Migration from existing data Importing from Postgres? See [Postgres ingest](https://originchaindb.com/docs/integrations/postgres-ingest). For CSV / JSON / NDJSON dumps, see the bulk-load section in [Insert → bulk](https://originchaindb.com/docs/insert#bulk). Connectors for DynamoDB and MongoDB are on the roadmap; in the meantime you can stream your dump through the standard bulk insert. --- # Elasticsearch-compatible API - OriginChainDB Canonical source: https://originchaindb.com/docs/elasticsearch Sitemap last modified: 2026-09-13T20:18:50.000Z elasticsearch # The Elasticsearch-compatible API OriginChainDB answers the Elasticsearch REST API and Query DSL over the same store your SQL and HTTP calls see. Point the official client at your instance and the requests you already write keep working — but the documents live in your database, so a write is searchable the moment it commits and your row-level security and column masking apply to the search itself. [Quickstart Connect a client and run your first search, end to end, in about five minutes.](https://originchaindb.com/docs/elasticsearch/quickstart)[Indexing Create an index, and decide what each field does - searchable text, exact keyword, or both.](https://originchaindb.com/docs/elasticsearch/indexing)[Search The Query DSL: match, bool, term, range, fuzzy, sorting and paging.](https://originchaindb.com/docs/elasticsearch/search)[Aggregations Terms, metrics, date histograms, named filter buckets and top hits.](https://originchaindb.com/docs/elasticsearch/aggregations)[Suggest Did-you-mean over the words actually indexed in a field.](https://originchaindb.com/docs/elasticsearch/suggest)[Bulk + concurrency Load many documents fast, and guard a write so it lands only if nobody changed the document first.](https://originchaindb.com/docs/elasticsearch/bulk)[By industry Worked examples: product catalog, application logs, transactions, support tickets.](https://originchaindb.com/docs/elasticsearch/industries)[All examples Sixteen copy-paste recipes, each one request and its response.](https://originchaindb.com/docs/examples/elasticsearch) ## What you point at it Any Elasticsearch 7.x client, or plain HTTP. The cluster reports version 7.14.2, which clears the product check the official clients run on connect. The base path is your tenant's ES endpoint; every index lives under it. ``` POST /v1/tenants/:tenant/es/:index/_search PUT /v1/tenants/:tenant/es/:index/_doc/:id POST /v1/tenants/:tenant/es/_bulk ``` ## What makes it different from a search cluster - No sync, no lag. A document is written by the same transaction that writes the row, so it is searchable the instant the call returns. There is no refresh interval to wait on and no connector to keep running. - Your policies apply to the search. Row-level security and column masking are enforced on the query itself, not just on the documents it returns. A caller who cannot see a row will not find it, and a query against a masked column is refused rather than answered. - One store. The same data answers SQL, the HTTP API, vector search and this. Nothing is copied anywhere. ## What it does not do This is a drop-in for the API real clients and dashboards exercise, not the entire Elasticsearch surface. Anything not answered faithfully is refused with an explicit error — it never returns a wrong or silently partial result. Each page below states its own limits where you meet them; the short list: | Area | Answered | Refused, and why | | --- | --- | --- | | Aggregations | terms, metrics, range, histogram, date_histogram, percentiles, top_hits, filters, sub-aggregations, pipelines | extended_stats, composite, nested have no executor yet; missing on a metric aggregation would fill in a value you did not choose | | Suggesters | the term suggester, with suggest_mode | phrase and completion have no executor; popular ranks by a document frequency this API does not report | | Paging | search_after within a 10,000-hit window | Point-in-Time and scroll hold a snapshot this engine does not take | | Scoring | function_score, per-field boosts | script_score and rescore would run arbitrary script | | Cluster management | index create, delete, mapping, _field_caps | index lifecycle, snapshots and templates are managed by OriginChainDB, not over this API | Vector and hybrid search live on their own surface — see [querying](https://originchaindb.com/docs/query). Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Elasticsearch aggregations on OriginChainDB - terms, metrics Canonical source: https://originchaindb.com/docs/elasticsearch/aggregations Sitemap last modified: 2026-09-13T20:18:50.000Z elasticsearch · aggregations # Aggregations Aggregations run over the whole match set, not the page of hits you asked for. Set size: 0 when you want the summary and not the documents. ## Group by a field A terms aggregation needs a field it can group on: keyword, numeric, date or boolean. On a text field use the .keyword sub-field — see [indexing](https://originchaindb.com/docs/elasticsearch/indexing). ``` await es.search({ index: 'shop.products', size: 0, aggs: { by_brand: { terms: { field: 'brand' } } } }) ``` ## Metrics avg, sum, min, max, value_count, stats and percentiles are answered, on their own or inside a bucket. ``` aggs: { avg_price: { avg: { field: 'price' } }, spread: { percentiles: { field: 'price' } }, by_brand: { terms: { field: 'brand' }, aggs: { avg_price: { avg: { field: 'price' } } } } } ``` missing is refused, not filled in missing on a metric aggregation asks us to substitute a value for documents that do not have the field. We refuse it rather than fold in a number you did not choose, because the result would look like data and be an assumption. ## Bucket by time A date_histogram groups documents into calendar buckets. The field can hold an ISO-8601 string or epoch milliseconds — both bucket the same way. Buckets are UTC and calendar-aligned, so pass calendar_interval. ``` aggs: { per_day: { date_histogram: { field: 'placed_at', calendar_interval: 'day', // minute hour day week month quarter year time_zone: 'UTC' } } } ``` fixed_interval is refused: an arbitrary fixed width has no calendar-aligned equivalent here, and approximating it would put your documents in the wrong buckets. ## Named filter buckets One bucket per named query, counting the documents that match both your search and that filter, with sub-aggregations over exactly those documents. other_bucket_key adds a bucket for everything that matched none of them. ``` aggs: { by_price: { filters: { other_bucket_key: 'mid_range', filters: { clearance: { range: { price: { lt: 50 } } }, premium: { range: { price: { gte: 200 } } } } }, aggs: { avg_price: { avg: { field: 'price' } } } } } ``` send filters in its own request A filters aggregation is answered with one search per filter, which is a different shape of work from the single pass every other aggregation takes. Combining it with other top-level aggregations in one request is refused, so that a request that looks like one search cannot quietly become several. ## The documents behind a bucket top_hits returns the best matching documents with their _source. It works at the top level, or nested so each bucket carries its own examples. ``` // three best matches overall aggs: { best: { top_hits: { size: 3 } } } // one example per brand, most expensive first aggs: { by_brand: { terms: { field: 'brand' }, aggs: { top: { top_hits: { size: 1, sort: [{ price: 'desc' }] } } } } } ``` ## Pipelines cumulative_sum and derivative post-process the buckets of the aggregation they sit in — a running total or a rate of change over a date histogram. ``` aggs: { per_day: { date_histogram: { field: 'placed_at', calendar_interval: 'day' }, aggs: { revenue: { sum: { field: 'total' } }, running: { cumulative_sum: { buckets_path: 'revenue' } } } } } ``` ## What is not answered - extended_stats, composite and nested have no executor yet and are refused rather than partially answered. - An aggregation that names no field is refused — except a bare top_hits, which legitimately needs none. - A field that was never indexed comes back as an empty aggregation with your hits intact, not as a failed search. [Correct the query first Did-you-mean over the words actually indexed.](https://originchaindb.com/docs/elasticsearch/suggest)[See it applied Catalog facets, log volume over time, transaction summaries.](https://originchaindb.com/docs/elasticsearch/industries) Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Elasticsearch _bulk ingest on OriginChainDB - if_seq_no Canonical source: https://originchaindb.com/docs/elasticsearch/bulk Sitemap last modified: 2026-09-13T20:18:50.000Z elasticsearch · bulk # Bulk ingest and safe concurrent writes _bulk returns HTTP 200 even when individual items fail — so read errors and each item’s own status. To stop two writers overwriting each other, carry if_seq_no and if_primary_term on an index action: the write lands only if the document is still the version you read. ## Load many documents Newline-delimited actions, exactly as a cluster takes them. An index that does not exist yet is created by the first write. ``` await es.bulk({ operations: [ { index: { _index: 'shop.products', _id: 'sku-1207' } }, { name: 'Trail 24', brand: 'Aero', price: 89 }, { index: { _index: 'shop.products', _id: 'sku-3355' } }, { name: 'City Runner', brand: 'Metro', price: 72 } ]}) ``` ## Read the result per item A bulk request is 200 even when some items fail; errors tells you whether any did, and each item carries its own status. Check it — a failed item in a successful request is the classic silent data-loss bug. ``` const res = await es.bulk({ operations }) if (res.errors) { for (const item of res.items) { const op = item.index ?? item.update ?? item.delete if (op.status >= 300) console.error(op._id, op.status, op.error?.type) } } ``` ## Guard a write against a concurrent one Carry if_seq_no and if_primary_term on an action and the write lands only if the document is still the version you read. A stale precondition comes back as a per-item 409 and every other item in the batch still applies, so the loser retries one document rather than the whole load. ``` // read the document you are about to change const cur = await es.get({ index: 'shop.products', id: 'sku-8842' }) await es.bulk({ operations: [ { index: { _index: 'shop.products', _id: 'sku-8842', if_seq_no: cur._version - 1, if_primary_term: 1 } }, { ...cur._source, price: 129 } ]}) // stale -> items[0].index.status === 409 // version_conflict_engine_exception ``` where preconditions apply On index actions. A create is already its own precondition (the document must not exist), so it does not take a version. A conditional update or delete is refused: an update is a read-modify-rewrite, and asserting the version at the rewrite is a different guarantee from the one your read saw. External versioning is refused too — it has no faithful mapping onto an exact compare-and-set. ## Partial updates update merges: the fields you send are written and every other field is preserved. ``` await es.bulk({ operations: [ { update: { _index: 'shop.products', _id: 'sku-8842' } }, { doc: { price: 139 } } // name, brand, ... untouched ]}) ``` ## Throughput Batches are written through the engine's chunked ingest path rather than one document at a time. Send documents in batches of a few hundred to a few thousand; a single enormous request is not faster and costs you the ability to retry a small piece of it. [Get the mapping right first A backfill into the wrong field type means reindexing it.](https://originchaindb.com/docs/elasticsearch/indexing)[Check what landed Count and group the data you just loaded.](https://originchaindb.com/docs/elasticsearch/aggregations) Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Create an Elasticsearch index and map fields - OriginChainDB Canonical source: https://originchaindb.com/docs/elasticsearch/indexing Sitemap last modified: 2026-09-13T20:18:50.000Z elasticsearch · indexing # Create an index, and choose what each field does An index is a table in your database, and its mapping decides what each field can do: full-text searchable, exact for grouping and sorting, or both. Declare a column text to search inside it, keyword to group and sort on it, and a multi-field when you need both. ## The short answer To make one column full-text searchable, declare it text when you create the index: ``` await es.indices.create({ index: 'shop.products', mappings: { properties: { description: { type: 'text' } // full-text searchable } } }) ``` That field now has a full-text index behind it: analyzed into terms, ranked by BM25, and reachable with match, match_phrase and fuzzy search. Everything below is how to choose when the answer is not that simple. ## What each type gives you | Type | Full-text search | Group / sort / aggregate | Use it for | | --- | --- | --- | --- | | text | yes, analyzed | no | prose: descriptions, messages, notes, titles you search inside | | keyword | exact value only | yes | identifiers, tags, statuses, brands, e-mail addresses | | long / double | no | yes | prices, counts, latencies, scores | | date | no | yes | timestamps, in ISO-8601 or epoch milliseconds | | boolean | no | yes | flags | the one that trips people up A text field cannot be grouped or sorted, and a keyword field cannot be searched inside. That is Elasticsearch's rule, not ours, and it is why the next section exists. ## When you need both on one column A product name you want to search inside and group by exactly needs two views of the same value. Declare a multi-field: the base is analyzed text, the sub-field is the exact keyword. ``` mappings: { properties: { name: { type: 'text', // match 'carbon marathon' fields: { keyword: { type: 'keyword' } } // group by the exact name } } } ``` Then search the base and aggregate the sub-field: ``` query: { match: { name: 'marathon' } } // searches the text aggs: { by_name: { terms: { field: 'name.keyword' } } } // groups the exact value ``` ## If you skip the mapping entirely You do not have to create the index first. Write a document to a name that does not exist and the index is created for you, with dynamic mapping: every string gets both a full-text index and a .keyword twin, numbers become numeric, and anything that parses as a date becomes a date. ``` await es.index({ index: 'shop.reviews', // does not exist yet - created by this write id: 'r-1', document: { body: 'fits well, runs small', stars: 4 } }) ``` That is the fastest way to start and the right default for logs and events. Declare a mapping when you want to be deliberate: to pick an analyzer, to keep a field out of the index, or to stop a string you never search from carrying a full-text index it does not need. ## Adding a field to an index that already exists Same call Elasticsearch uses. New fields are added; existing ones keep their type. ``` await es.indices.putMapping({ index: 'shop.products', properties: { supplier: { type: 'keyword' } } }) ``` a field's type is fixed once it exists Changing text to keyword on a live field would make every document already written disagree with the mapping. Create a new index with the mapping you want and _reindex into it — the same move you would make on a real cluster. ## Checking what a field actually is _field_caps reads the stored mapping and tells you, per field, what it is and whether it can be searched or aggregated. This is the call dashboards make on connect, and the fastest way to find out why an aggregation came back empty. ``` await es.fieldCaps({ index: 'shop.products', fields: '*' }) // name -> text searchable: true, aggregatable: false // brand -> keyword searchable: true, aggregatable: true // price -> long searchable: true, aggregatable: true ``` ## Analyzers, when the language matters A text field can declare an analyzer so that searching for running finds run. Set it per field in the mapping; the same analyzer is used when the document is indexed and when the query is parsed, so the two always agree. ``` mappings: { properties: { description: { type: 'text', analyzer: 'english' } } } ``` An analyzer we cannot reproduce faithfully is refused when you create the index, rather than accepted and quietly ignored. ## Index names fold Index names are matched with case and separators folded, so Shop-Products, shop_products and shop.products are one index, not three. If a name would land on an index another name already owns, the create is refused with the owner named, rather than quietly writing into it. ``` // shop.products exists await es.indices.create({ index: 'Shop-Products' }) // -> 400 already exists: [Shop-Products] differs from it only in // case or separators and would alias its data ``` Pick one spelling and keep it. The refusal exists so a typo cannot silently merge two datasets. ## Deleting an index ``` await es.indices.delete({ index: 'shop.reviews' }) ``` This removes the index and its documents. A delete of a name that folds onto an index another name owns is refused, for the same reason a create is. [Now search it match, bool, term, range, fuzzy, sorting and paging.](https://originchaindb.com/docs/elasticsearch/search)[Load it properly Bulk ingest, and guarding a write against a concurrent one.](https://originchaindb.com/docs/elasticsearch/bulk) Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Elasticsearch examples by industry - OriginChainDB Canonical source: https://originchaindb.com/docs/elasticsearch/industries Sitemap last modified: 2026-09-13T20:18:50.000Z elasticsearch · by industry # Worked examples by industry Four models, each starting with the mapping — because the mapping decides what you can ask later — then the queries that pay for it: a product catalog with facets, an application log stream, a payment ledger and a support ticket desk. ## Retail: a product catalog with facets A search box over product text, plus the filters down the side of the page: brand, price band, availability. The trap here is wanting to search a name and group by it, which needs both views of the same field. ### The mapping Name and description are prose, so they are text. Brand and colour are picked from a list, so they are keyword. Name also carries a .keyword twin so it can be grouped exactly. ``` mappings: { properties: { name: { type: 'text', fields: { keyword: { type: 'keyword' } } }, description: { type: 'text', analyzer: 'english' }, brand: { type: 'keyword' }, colour: { type: 'keyword' }, price: { type: 'long' }, in_stock: { type: 'boolean' }, added: { type: 'date' } } } ``` ### Search plus the facet counts, in one request The search ranks by text relevance while the filters narrow it, and the aggregations count what is left so the sidebar numbers always match the results. ``` await es.search({ index: 'shop.products', size: 20, query: { bool: { must: [{ match: { description: 'waterproof running' } }], filter: [{ term: { in_stock: true } }] } }, aggs: { by_brand: { terms: { field: 'brand' } }, by_colour: { terms: { field: 'colour' } }, avg_price: { avg: { field: 'price' } } } }) ``` ### Price bands, and a sample product per band Bands are not a field, they are a question, so they are named filters. Send this one on its own: a filters aggregation is answered with one search per filter. ``` aggs: { bands: { filters: { other_bucket_key: 'mid', filters: { budget: { range: { price: { lt: 50 } } }, premium: { range: { price: { gte: 200 } } } } }, aggs: { example: { top_hits: { size: 1, _source: ['name','price'] } } } } } ``` when a shopper mistypes Run the search and a [suggester](https://originchaindb.com/docs/elasticsearch/suggest) in the same request. If the search returns nothing, show the correction instead of an empty page. ## Software: an application log stream High write volume, mostly time-ranged reads, and one question repeated all day: what broke, where, and when did it start. ### The mapping The message is searched, so text. Service, level and host are grouped and filtered, so keyword. The timestamp drives every chart, so it is a date — an ISO string or epoch milliseconds, either works. ``` mappings: { properties: { message: { type: 'text' }, service: { type: 'keyword' }, level: { type: 'keyword' }, host: { type: 'keyword' }, latency_ms:{ type: 'long' }, ts: { type: 'date' } } } ``` ### Errors per hour for one service A date histogram over the filtered set. Buckets are UTC and calendar-aligned. ``` await es.search({ index: 'ops.logs', size: 0, query: { bool: { filter: [ { term: { service: 'checkout' } }, { term: { level: 'error' } }, { range: { ts: { gte: '2026-09-05T00:00:00Z' } } } ] } }, aggs: { per_hour: { date_histogram: { field: 'ts', calendar_interval: 'hour', time_zone: 'UTC' } } } }) ``` ### The worst offenders, with an example line each Group by service, then pull the slowest request in each group so the chart has something to click into. ``` aggs: { by_service: { terms: { field: 'service' }, aggs: { p95: { percentiles: { field: 'latency_ms' } }, worst:{ top_hits: { size: 1, sort: [{ latency_ms: 'desc' }], _source: ['message','host','ts'] } } } } } ``` ingest shape Send logs with _bulk in batches of a few hundred to a few thousand, and check errors on the response — a failed item inside a successful request is the classic way to lose data quietly. See [bulk ingest](https://originchaindb.com/docs/elasticsearch/bulk). ## Financial services: a payment ledger Amounts and dates you summarise, counterparties you look up exactly, and a free-text reference people actually search. The distinguishing requirement is that not everyone may see every row. ### The mapping The reference is prose. Everything you group or compare is keyword, long or date. Money is stored in minor units as an integer, so totals never drift. ``` mappings: { properties: { reference: { type: 'text' }, counterparty:{ type: 'keyword' }, status: { type: 'keyword' }, currency: { type: 'keyword' }, amount_minor:{ type: 'long' }, // 12345 = 123.45 booked_at: { type: 'date' } } } ``` ### Daily totals by currency A date histogram with a nested terms aggregation and a sum, which is one request rather than a report job. ``` await es.search({ index: 'fin.payments', size: 0, query: { bool: { filter: [{ term: { status: 'settled' } }] } }, aggs: { per_day: { date_histogram: { field: 'booked_at', calendar_interval: 'day' }, aggs: { by_ccy: { terms: { field: 'currency' }, aggs: { total: { sum: { field: 'amount_minor' } } } } } } } }) ``` who can see what is not your query's problem Row-level security and column masking are enforced on the search itself, so a caller restricted to one desk finds only that desk's rows — through any query, including aggregations. A search against a masked column is refused rather than answered, because a match set could otherwise be used to reconstruct the value the mask hides. You do not add a filter for this, and you cannot forget to. ## Support: a ticket desk Agents search what customers wrote, managers want queue health, and both want it in the same shape. ### The mapping Subject and body are searched. Queue, status and priority are grouped. First response time is a number you average. ``` mappings: { properties: { subject: { type: 'text' }, body: { type: 'text', analyzer: 'english' }, queue: { type: 'keyword' }, status: { type: 'keyword' }, priority: { type: 'keyword' }, first_reply_mins: { type: 'long' }, opened_at: { type: 'date' } } } ``` ### Open tickets mentioning a problem, with queue health beside them One request answers the agent's search and the manager's dashboard, over the same match set. ``` await es.search({ index: 'support.tickets', size: 20, query: { bool: { must: [{ match: { body: 'refund not received' } }], filter: [{ term: { status: 'open' } }] } }, aggs: { by_queue: { terms: { field: 'queue' }, aggs: { avg_first_reply: { avg: { field: 'first_reply_mins' } } } } } }) ``` ### Two agents editing one ticket Read the ticket, then write it back conditionally. If someone else saved first, you get a conflict instead of silently overwriting their edit. ``` const cur = await es.get({ index: 'support.tickets', id }) await es.bulk({ operations: [ { index: { _index: 'support.tickets', _id: id, if_seq_no: cur._version - 1, if_primary_term: 1 } }, { ...cur._source, status: 'resolved' } ]}) // someone else saved first -> items[0].index.status === 409 ``` ## What these four have in common - The mapping is the design. Every question you can ask later is decided by which fields are text and which are keyword. See [indexing](https://originchaindb.com/docs/elasticsearch/indexing). - Search and summary in one request. Hits and aggregations come from the same match set, so a dashboard cannot disagree with the list beneath it. - No copy to keep in sync. These indexes are your data, not a projection of it, so a write is searchable when it commits. [Start from zero Connect a client and run the first search.](https://originchaindb.com/docs/elasticsearch/quickstart)[All examples Sixteen single-purpose recipes with their responses.](https://originchaindb.com/docs/examples/elasticsearch) Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Elasticsearch quickstart - connect a client - OriginChainDB Canonical source: https://originchaindb.com/docs/elasticsearch/quickstart Sitemap last modified: 2026-09-23T08:33:23.000Z elasticsearch · quickstart # Connect and run your first search Connect an Elasticsearch client, index sample documents, and run a search against OriginChainDB. ## Before you start You need two things. The endpoint is your instance URL with /v1/tenants//es on the end — there is nothing to enable, every instance answers it. The API key comes from the console. The endpoint already carries your tenant, so a client treats it like any cluster URL. ``` export OC_ES_URL='https:///v1/tenants//es' export OC_API_KEY='' ``` ## The five steps 1. 1. Install a client Any Elasticsearch 7.x client works. The Node client is used here; the same requests work from Python, Java, Go, or plain curl. ``` npm install @elastic/elasticsearch@7 ``` 2. 2. Point it at your instance Pass the endpoint as the node and your API key as the auth. info() clears the product check and reports the version. ``` import { Client } from '@elastic/elasticsearch' const es = new Client({ node: process.env.OC_ES_URL, auth: { apiKey: process.env.OC_API_KEY } }) await es.info() // version 7.14.2 ``` 3. 3. Create an index with a mapping Declare your fields. text is analyzed for full-text search; keyword is exact and can be grouped and sorted. [Indexing](https://originchaindb.com/docs/elasticsearch/indexing) explains how to choose, and what happens if you skip this step. ``` await es.indices.create({ index: 'shop.products', mappings: { properties: { name: { type: 'text' }, brand: { type: 'keyword' }, price: { type: 'long' }, added: { type: 'date' } } } }) ``` 4. 4. Add a document It is searchable the moment this returns. There is no refresh to wait for. ``` await es.index({ index: 'shop.products', id: 'sku-8842', document: { name: 'Carbon Marathon', brand: 'Aero', price: 149, added: '2026-09-01T00:00:00Z' } }) ``` 5. 5. Search it The Query DSL you already write. This one matches the analyzed text and filters on the exact brand. ``` const res = await es.search({ index: 'shop.products', query: { bool: { must: [{ match: { name: 'marathon' } }], filter: [{ term: { brand: 'Aero' } }] } } }) // res.hits.total.value === 1 // res.hits.hits[0]._source.name === 'Carbon Marathon' ``` ## Check it from the shell The same thing without a client, so you can prove connectivity before wiring an app: ``` curl "$OC_ES_URL/shop.products/_search" \ -H "Authorization: Bearer $OC_API_KEY" \ -H "Content-Type: application/json" \ -d '{"query":{"match":{"name":"marathon"}}}' ``` if the first call fails A 401 means the API key is not being sent — check the auth option, or the Authorization: Bearer header if you are using curl. A 404 on a search means the index does not exist yet: create it, or just write a document to it and it will be created for you. A client that refuses to talk at all is usually pinned to the 8.x line — use a 7.x client. ## Where to go next [Choose what each field does Text, keyword, or both - and how to make one existing column searchable.](https://originchaindb.com/docs/elasticsearch/indexing)[Load real data Bulk ingest, and how to guard a write against a concurrent one.](https://originchaindb.com/docs/elasticsearch/bulk)[Summarise it Group, count, average, bucket by day, or pull the top documents per bucket.](https://originchaindb.com/docs/elasticsearch/aggregations)[See it modelled Worked examples for a catalog, a log stream, transactions and tickets.](https://originchaindb.com/docs/elasticsearch/industries) Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Elasticsearch Query DSL on OriginChainDB - match, bool, term Canonical source: https://originchaindb.com/docs/elasticsearch/search Sitemap last modified: 2026-09-13T20:18:50.000Z elasticsearch · search # Search with the Query DSL The Query DSL you already write works unchanged: match and match_phrase for analyzed text, term and terms for exact values, plus range, bool and fuzziness. Deep paging is search_after within a 10,000-hit window, and anything this engine cannot answer faithfully is refused rather than approximated. ## Full-text: match and match_phrase match analyzes your terms the same way the field was analyzed and ranks by BM25. match_phrase keeps the words adjacent. ``` query: { match: { description: 'waterproof jacket' } } // only this exact phrase, in this order query: { match_phrase: { description: 'waterproof jacket' } } ``` ## Exact: term and terms Use these on keyword, numeric, date and boolean fields. On an analyzed text field a term query looks for a single analyzed token, which is rarely what you want. ``` query: { term: { brand: 'Aero' } } query: { terms: { brand: ['Aero', 'Metro'] } } ``` ## Ranges ``` query: { range: { price: { gte: 50, lt: 200 } } } query: { range: { added: { gte: '2026-09-01T00:00:00Z' } } } ``` ## Combining: bool must contributes to the score, filter does not (and is the one to use for exact predicates), should boosts, must_not excludes. ``` query: { bool: { must: [{ match: { name: 'marathon' } }], filter: [{ term: { brand: 'Aero' } }, { range: { price: { lte: 200 } } }], should: [{ match: { description: 'carbon' } }], must_not: [{ term: { discontinued: true } }] } } ``` ## Typo tolerance Fuzziness walks the words actually indexed in the field. For a search box, pair it with [the suggester](https://originchaindb.com/docs/elasticsearch/suggest) so you can show what you corrected to. ``` query: { match: { name: { query: 'marathn', fuzziness: 'AUTO' } } } ``` ## Sorting, paging and trimming the response ``` sort: [{ price: 'desc' }, '_score'], from: 0, size: 20, _source: ['name', 'price'] // only these fields come back ``` For deep paging use search_after with the sort values of the last hit, within a 10,000-hit window. Point-in-Time and scroll are not offered: both hold a snapshot this engine does not take, and pretending otherwise would hand you a page that quietly drifts. ## Reading one document When you know the id, skip the search. The version that comes back is the one the write reported, so you can carry it into a conditional write. ``` await es.get({ index: 'shop.products', id: 'sku-8842' }) await es.exists({ index: 'shop.products', id: 'sku-8842' }) // true / false await es.mget({ index: 'shop.products', body: { ids: ['a','b'] } }) ``` ## What a missing index does A search or count against an index that does not exist answers 404 index_not_found_exception, the same as a real cluster. An index that exists but holds nothing answers 200 with zero hits. Those are different answers on purpose: a client polling for an index can tell absent from empty. ## Security is part of the query Row-level security and column masking are enforced on the search itself. A caller who cannot see a row will not find it through any query, and a query that searches a masked column is refused rather than answered — otherwise the match set could be used to reconstruct the value the mask hides. [Summarise results Terms, metrics, date histograms, filter buckets and top hits.](https://originchaindb.com/docs/elasticsearch/aggregations)[Copy-paste recipes Sixteen worked requests with their responses.](https://originchaindb.com/docs/examples/elasticsearch) Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Elasticsearch term suggester on OriginChainDB - did-you-mean Canonical source: https://originchaindb.com/docs/elasticsearch/suggest Sitemap last modified: 2026-09-13T20:18:50.000Z elasticsearch · suggest # Did you mean The term suggester corrects a word against the vocabulary actually indexed in a field. Use it when a search returns nothing, to offer the spelling that would have worked. ## Ask for a correction ``` await es.search({ index: 'shop.products', size: 0, suggest: { did_you_mean: { text: 'marathn', term: { field: 'name' } } } }) // suggest.did_you_mean[0] // text: 'marathn', offset: 0, length: 7 // options: [ { text: 'marathon', score: 0.857 } ] ``` Each entry is one word of your input, with its offset and length so you can splice a correction back into the original string. Options come back best first; the score is one minus the edit distance over the word length. ## Correct every word, or only unknown ones suggest_mode defaults to missing: a word the index already holds gets no options, which is what a search box wants. always returns near neighbours even for a word that exists, which is what you want for query expansion. ``` term: { field: 'name', suggest_mode: 'always', size: 3 } ``` ## A search box, end to end Run the search; if it returns nothing, ask for a correction and offer it. ``` const res = await es.search({ index: 'shop.products', query: { match: { name: q } }, suggest: { fix: { text: q, term: { field: 'name' } } } }) if (res.hits.total.value === 0) { const best = res.suggest.fix[0]?.options[0]?.text if (best) console.log(`Did you mean ${best}?`) } ``` ## What it does not return no document frequency Elasticsearch puts a freq on every option. That number lives somewhere this API cannot reach, so the field is omitted rather than invented. If you rank suggestions yourself, rank on score. For the same reason suggest_mode: popular, which is defined by frequency, is refused. The phrase and completion suggesters are refused by name — there is no executor for either, and a silent empty result would read as "no suggestions" rather than "not supported". [Fuzzy matching Typo tolerance inside the query itself, rather than as a correction.](https://originchaindb.com/docs/elasticsearch/search) Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # Error reference - OriginChainDB HTTP error codes Canonical source: https://originchaindb.com/docs/errors Sitemap last modified: 2026-09-13T20:18:50.000Z reference · errors # Error reference Every non-2xx response from the OriginChainDB API has a JSON body with a stable `error` code, a human-readable `message`, and a `retry` flag that tells you whether the SDK should retry automatically. ``` // Every error response looks like this: { "error": "schema_not_found", "message": "no schema registered under 'shop.orders'", "retry": false } ``` Switch on the `error` code in your application, not on the message - codes are stable across versions; messages are tuned for humans and may change. ## Client errors (4xx). Your request was malformed, unauthorized, or referred to something that doesn't exist. Fix the request; don't retry blindly. | Code | HTTP | What it means | How to fix | | --- | --- | --- | --- | | invalid_body | 400 | Request body couldn't be parsed. | Check Content-Type matches the body (application/json for JSON, text/plain for TOML schemas, application/x-ndjson for streaming inserts). | | type_mismatch | 400 | A column value doesn't match the declared column type. | Check the error message - it names the offending field. Common: sending a string for an i64 column, or a number for a str. | | missing_required | 400 | A column declared `required = true` was omitted from a row write. | Send a value for every required column. Schema reference shows which columns are required. | | fk_violation | 400 | A foreign key references a row that doesn't exist. | Insert the target row first, or remove the foreign-key constraint if you really want orphan references. | | check_violation | 400 | A row failed a CHECK constraint. | Inspect the row's values against the constraint expression. The error message names the failing constraint. | | dim_mismatch | 400 | A vector's length doesn't match the table's locked dim. | Every vector in a table must have the same dim. The first put sets it. | | metric_mismatch | 400 | A vector put used a different metric than the table's locked metric. | Pick one of cosine / dot / l2 / manhattan and use it for every put + topk on the table. | | unsupported_sql | 400 | The SQL statement contains a construct that isn't yet implemented. | See SQL reference for the supported surface. Common: ORDER BY, HAVING, OR in WHERE, window functions, CTEs. | | validation_failed | 400 | Generic catch-all for malformed input. | Read the message field - it always names the specific problem. | | unauthorized | 401 | Missing, malformed, or unknown bearer token. | Check your Authorization header. See Auth reference. | | forbidden | 403 | Token valid but doesn't have access to this resource. | Usually means the token belongs to a different instance. Check endpoint hostname matches the token. | | addon_required | 402 | Endpoint requires a paid add-on that isn't enabled. | Enable the add-on in the dashboard, or call the alternative endpoint that doesn't need it. | | schema_not_found | 404 | No schema is registered under the given name. | Register the schema first - POST /v1/tenants/:t/schemas with the TOML manifest. | | row_not_found | 404 | GET /rows/:schema/:pk for a primary key that doesn't exist. | Verify the pk. If the row was deleted, this is the expected response. | | conflict | 409 | ?expect=insert was set and the row already exists. | Either drop the expect param (will overwrite) or update via the normal write path. | | rate_limited | 429 | Per-key request budget exceeded. | Honor the Retry-After header. See Rate limits reference. | ## Server errors (5xx). Something went wrong on our side. Retry with exponential backoff. The SDKs do this automatically (up to 3 attempts by default). | Code | HTTP | What it means | How to fix | | --- | --- | --- | --- | | internal_error | 500 | Server-side bug - the engine hit something it didn't expect. | Retry once. If it persists, file a support ticket with the X-OC-Trace-Id from the response headers. | | service_unavailable | 503 | The engine is temporarily unable to serve the request (often during failover or migration). | Retry with exponential backoff. The SDKs do this automatically up to 3 retries. | | gateway_timeout | 504 | Upstream timeout (replication ack didn't arrive in budget). | Retry. If persistent, check the instance's replication health in the dashboard. | ## Retry strategy. The SDKs already implement the right policy. If you're using raw HTTP, follow the same pattern: - 429: read the `Retry-After` header (seconds), wait that long, retry. - 500, 502, 503, 504: exponential backoff (0.25s, 0.5s, 1s, 2s, capped at 4s). Maximum 3 retries. - All 4xx except 429: do not retry. The request is malformed; retrying won't help. - Idempotency: the SDKs auto-attach an `Idempotency-Key` header on every mutating call, so retries don't double-insert. For implementation details, see [Rate limits](https://originchaindb.com/docs/rate-limits). --- # Examples - 73 copy-paste query recipes - OriginChainDB Canonical source: https://originchaindb.com/docs/examples Sitemap last modified: 2026-09-13T20:18:50.000Z examples # Examples There are 73 worked examples here, spread across 8 sections, covering every query shape OriginChainDB supports. Each one is its own page carrying the schema and seed data it needs, side-by-side cURL / Python / TypeScript / Go, the response shape, and the mistakes that break it. Pick a section to see its index. The count on each card is the number of examples in it. [SQL examples 13 SELECT, WHERE filters, JOINs, GROUP BY + aggregates, writes via /sql, plan inspection. Plus a few roadmap features so you know what's coming.](https://originchaindb.com/docs/examples/sql)[Vector examples 7 Single insert, insert with metadata, top-k by each metric, fast vs high_recall mode, filtered top-k. Working request and response bodies.](https://originchaindb.com/docs/examples/vector)[Full-text examples 6 Index a doc, boolean AND, exact phrase, ranked BM25, and multi-field weighted merge. Side-by-side cURL and Python.](https://originchaindb.com/docs/examples/fts)[Elasticsearch examples 10 Point the @elastic client at OriginChainDB: create an index, index and _bulk-load docs, match / bool queries, aggregations, search_after, and delete_by_query.](https://originchaindb.com/docs/examples/elasticsearch)[Graph examples 7 Forward and reverse neighbors, depth-bounded BFS, reachability, weighted shortest path, all simple paths, PageRank. One example per page.](https://originchaindb.com/docs/examples/graph)[Natural-language examples 6 Simple count, schemas hint, show_plan, NL JOIN, cache-hit replay, error recovery. The full /ask round-trip with rows + cache metadata.](https://originchaindb.com/docs/examples/ask)[Atomic multi-shape writes 5 Recipes that save a row, its vector embedding, its full-text index, and its graph edge in one coordinated write.](https://originchaindb.com/docs/examples/atomic-multi-shape)[Errors 19 Every HTTP error code you can hit, with the request that triggers it and the canonical response body. Includes addon-required, rate-limit, and conflict cases.](https://originchaindb.com/docs/examples/errors) Looking for the per-shape reference instead of recipes? See [SQL](https://originchaindb.com/docs/sql), [vector](https://originchaindb.com/docs/vector), [full-text](https://originchaindb.com/docs/fts), [graph](https://originchaindb.com/docs/graph), and [natural language](https://originchaindb.com/docs/ask). For schema design, see [Schemas](https://originchaindb.com/docs/schemas). --- # Ask (natural-language) examples - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/ask Sitemap last modified: 2026-09-13T20:18:50.000Z examples · ask # Ask examples [← All examples](https://originchaindb.com/docs/examples) POST /v1/tenants/:t/ask takes plain English in the `"nl"` field, compiles it into a Plan tree against your registered schemas, and returns rows plus a cache flag. Six examples follow, each on its own page. 6 work today. Each one shows the request body the engine accepts, the JSON it returns, and the ways a question can go wrong. [1 Simple count question works today The minimum Ask call. Send a plain-English question, get rows and a cache flag back. →](https://originchaindb.com/docs/examples/ask/simple-count)[2 Narrow with a schemas hint works today Pass the schemas list when you know which table the question is about. Faster, more predictable plans. →](https://originchaindb.com/docs/examples/ask/schemas-hint)[3 Return the compiled Plan tree works today Set show_plan: true in the body to see the executor plan. Verify the index path before you put a question in a hot loop. →](https://originchaindb.com/docs/examples/ask/show-plan)[4 JOIN expressed in English works today The compiler walks the relations between schemas and produces a JOIN underneath. Be specific about which column you mean. →](https://originchaindb.com/docs/examples/ask/join-via-nl)[5 Plan-cache hit on repeat call works today Same question twice. The second call lands on cache: "hit" because the canonicalised question + schemas key already has a plan. →](https://originchaindb.com/docs/examples/ask/cache-hit)[6 When the compiler can't make a plan works today Gibberish in, structured error out. 400 with a no_plan_compiled body, a hint, and a suggested rephrase. →](https://originchaindb.com/docs/examples/ask/error-recovery) --- # Plan-cache hit on repeat call - OriginChainDB Ask example Canonical source: https://originchaindb.com/docs/examples/ask/cache-hit Sitemap last modified: 2026-09-13T20:18:50.000Z examples · ask · 5 / 6 # 5. Plan-cache hit on repeat call [← Ask examples](https://originchaindb.com/docs/examples/ask) what this does The first Ask call compiles a plan and stores it under a key derived from the canonicalised question and the `schemas` list. An identical second call finds that plan and skips compilation - the `cache` field in the response flips from `"miss"` to `"hit"`. when to use it - The Ask cache is automatic - you don't opt in. This example just shows the behaviour you can rely on. - Useful as a quick check: see `cache: "miss"` on every call? Your questions are varying in a way that's defeating the canonical form (extra punctuation, embedded values, etc.). - If you want a fresh compile on purpose, change the question or change the schemas list - both are part of the key. the schema ``` namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "email" ty = "str" ``` seed data ``` # Any handful of customers is fine - the cache hit doesn't depend on the count. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.customers/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "c_1", "email": "alice@example.com" }, { "id": "c_2", "email": "bob@example.com" } ]' ``` the request #### cURL ``` # Call the same question twice - watch the cache field flip. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "how many customers do I have?", "schemas": ["shop.customers"] }' # Same body, same headers - just send it again. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "how many customers do I have?", "schemas": ["shop.customers"] }' ``` #### Python ``` first = db.ask("how many customers do I have?", schemas=["shop.customers"]) second = db.ask("how many customers do I have?", schemas=["shop.customers"]) print(first["cache"]) # "miss" print(second["cache"]) # "hit" ``` #### TypeScript ``` const first = await db.ask("how many customers do I have?", { schemas: ["shop.customers"] }); const second = await db.ask("how many customers do I have?", { schemas: ["shop.customers"] }); console.log(first.cache); // "miss" console.log(second.cache); // "hit" ``` #### Go ``` first, _ := db.Ask(ctx, "how many customers do I have?", "shop.customers") second, _ := db.Ask(ctx, "how many customers do I have?", "shop.customers") fmt.Println(first.Cache) // "miss" fmt.Println(second.Cache) // "hit" ``` what you get back First call: ``` { "rows": [ { "count": 2 } ], "cache": "miss" } ``` Second call: ``` { "rows": [ { "count": 2 } ], "cache": "hit" } ``` The `cache` values are `"hit"`, `"miss"`, or `"skip"`. `"skip"` means the engine bypassed the cache - for example when the question depends on state that can't be safely cached. how it works - The cache key is `(canonicalised_nl, sorted_schemas_list)`. Case, whitespace, and trailing punctuation are normalised before hashing. - Only the plan tree is cached - rows are always read fresh from the row store, so the count reflects current state on a `"hit"`. - Schema migrations evict entries that reference affected schemas. After you alter `shop.customers`, the next call for that question re-plans. - The cache is per-tenant and durable across replicas - it rides the WAL stream that provisioned standbys tail, so a standby that is promoted keeps the entries it has already tailed. common mistakes - Interpolating values into the question. `"orders from customer c_1"` and `"orders from customer c_2"` are different cache entries. Use a parameterised question: `"orders for a given customer id"` and bind the value separately when supported, or fall back to `/v1/query` with a saved plan. - Re-ordering the schemas list. `["a","b"]` and `["b","a"]` hash to the same key (the engine sorts), but adding or removing a schema does change it. - Expecting stale rows on a hit. The plan is cached. The data is not. A `"hit"` still runs the executor against current rows. --- # no_plan_compiled - Ask error recovery - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/ask/error-recovery Sitemap last modified: 2026-09-13T20:18:50.000Z examples · ask · 6 / 6 # 6. When the compiler can't make a plan [← Ask examples](https://originchaindb.com/docs/examples/ask) what this does When the compiler cannot bind a question to your schemas it refuses to guess and returns a `400` with a structured body: a stable error code (`no_plan_compiled`), a human-readable hint explaining why, and a suggested rephrase that uses identifiers actually present in your schemas. when this happens - The question doesn't contain any recognisable table, column, or relation name. - The question is ambiguous in a way the compiler can't resolve - two equally-good schemas match and there's no hint to break the tie. - The question implies a feature that isn't supported (e.g. recursive CTEs, full-text search) - you'll get `no_plan_compiled` with a hint pointing at the unsupported feature. Show the suggested rephrase in your UI when you have one. It's grounded in the user's actual schemas, not boilerplate. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "ufnwiefwn wefjwe" }' ``` #### Python ``` try: db.ask("ufnwiefwn wefjwe") except OriginChainError as e: print(e.status) # 400 print(e.code) # "no_plan_compiled" print(e.hint) print(e.suggested_rephrase) ``` #### TypeScript ``` try { await db.ask("ufnwiefwn wefjwe"); } catch (e) { if (e.status === 400 && e.code === "no_plan_compiled") { console.log(e.hint); console.log(e.suggested_rephrase); } } ``` #### Go ``` _, err := db.Ask(ctx, "ufnwiefwn wefjwe") if err != nil { var apiErr *originchain.APIError if errors.As(err, &apiErr) && apiErr.Code == "no_plan_compiled" { fmt.Println(apiErr.Hint) fmt.Println(apiErr.SuggestedRephrase) } } ``` what you get back ``` HTTP/1.1 400 Bad Request { "error": "no_plan_compiled", "hint": "the question contained no recognisable identifiers (table, column, or relation name). nothing in the registered schemas matched any token.", "suggested_rephrase": "ask in terms of names that appear in your schemas - e.g. \"how many shop.customers do I have\"" } ``` `error` is stable - branch on it in your code. `hint` and `suggested_rephrase` are human-readable and may change wording between releases; show them, don't parse them. how it works - The compiler tokenises the question and tries to bind tokens to schema identifiers (namespace, table, column, relation name). - If too few tokens bind, the compiler refuses to guess. There is no silent fallback to "scan everything". - The hint explains the binding failure in human terms; the suggested rephrase uses identifiers from the visible catalog as anchor words. - A failed compile is not cached - retry safely with a better question. common mistakes - Ambiguous column references. "Total by month" with no date column anywhere in the visible schemas is unanchored - say which column to bucket by. - Sensitive data in questions. The LLM-fallback path (if configured) sees the full question. Don't put PII, secrets, or credentials in the `nl` field - bind those as parameters in `/v1/query` instead. - Retrying without rephrasing. The error is deterministic for a fixed schema set. Retrying the same question gets the same error. Either rephrase or change the schemas list (or registered schemas). --- # JOIN expressed in English - OriginChainDB Ask example Canonical source: https://originchaindb.com/docs/examples/ask/join-via-nl Sitemap last modified: 2026-09-13T20:18:50.000Z examples · ask · 4 / 6 # 4. JOIN expressed in English [← Ask examples](https://originchaindb.com/docs/examples/ask) what this does "Total order amount per customer email" needs a join: `orders` has the amount, `customers` has the email. The compiler sees the declared relation on `customer_id` and folds it into a hash join under an aggregate. when to use it - The question pulls fields from more than one table and you've declared the relation between them. - You want the answer phrased in terms of a human-readable field (email) rather than the underlying foreign key (customer_id). - You'd rather not maintain hand-rolled JOIN SQL for evolving reports. List both schemas in the hint. The compiler can only join across schemas that are in the visible catalog for this call. the schemas Register both. The `[[relations]]` block on `shop.orders` is what makes the join discoverable. ``` # Two related schemas. The relation lets the compiler walk customer_id # from orders to id on customers. namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "email" ty = "str" --- namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer_id" ty = "str" [[columns]] name = "amount" ty = "f64" [[relations]] name = "customer" columns = ["customer_id"] references = { schema = "shop.customers", columns = ["id"] } ``` seed data ``` # Two customers, three orders between them. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.customers/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "c_1", "email": "alice@example.com" }, { "id": "c_2", "email": "bob@example.com" } ]' curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "o_1", "customer_id": "c_1", "amount": 19.00 }, { "id": "o_2", "customer_id": "c_1", "amount": 43.40 }, { "id": "o_3", "customer_id": "c_2", "amount": 129.00 } ]' ``` the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "total order amount per customer email", "schemas": ["shop.orders", "shop.customers"] }' ``` #### Python ``` result = db.ask( "total order amount per customer email", schemas=["shop.orders", "shop.customers"], ) for row in result["rows"]: print(row["email"], row["total"]) ``` #### TypeScript ``` const result = await db.ask( "total order amount per customer email", { schemas: ["shop.orders", "shop.customers"] } ); for (const row of result.rows) { console.log(row.email, row.total); } ``` #### Go ``` result, err := db.Ask(ctx, "total order amount per customer email", "shop.orders", "shop.customers", ) if err != nil { /* handle */ } for _, row := range result.Rows { fmt.Println(row["email"], row["total"]) } ``` what you get back ``` { "rows": [ { "email": "alice@example.com", "total": 62.40 }, { "email": "bob@example.com", "total": 129.00 } ], "cache": "miss" } ``` One row per distinct email, with the summed amount. Order is by the underlying hash-aggregator's bucket order - if you need a specific order, add "sorted by total descending" to the question. how it works - The compiler reads both schemas and discovers the relation on `shop.orders.customer_id → shop.customers.id`. - The phrase "per customer email" pins the grouping column to `customers.email`, which forces a join. - The resulting plan is `Aggregate(sum amount) → HashJoin(orders.customer_id = customers.id) → Scan(orders) + Scan(customers)`. - Set `show_plan: true` to see which side of the join the planner picked as the build vs probe. common mistakes - Ambiguous "by customer". "Per customer" alone leaves the grouping column under-specified - the compiler may pick `customer_id`. Say "per customer email" or "per customer name" so the result is human-readable. - Missing the relation declaration. Without `[[relations]]` on one side, the compiler has no way to join the two schemas and returns `no_plan_compiled`. - Only listing one schema in the hint. The hint is an upper bound on what the compiler can see. If only `shop.orders` is in the list, customers is invisible and the join fails to compile. --- # Narrow with a schemas hint - OriginChainDB Ask example Canonical source: https://originchaindb.com/docs/examples/ask/schemas-hint Sitemap last modified: 2026-09-13T20:18:50.000Z examples · ask · 2 / 6 # 2. Narrow with a schemas hint [← Ask examples](https://originchaindb.com/docs/examples/ask) what this does A `"schemas"` array in the Ask body restricts the compiler to the schemas you list, so it never has to disambiguate against schemas in other namespaces. The endpoint is the same POST /v1/tenants/:t/ask as the simple count. when to use it - You know which table the question is about - typical for app code where the surface is constrained. - You have multiple namespaces in the same tenant and the same noun appears in more than one (`crm.customers` vs `shop.customers`). - You're in a hot path and want compile cost to be a function of N schemas you control, not the whole tenant catalog. The hint is included in the plan-cache key. Two calls with the same question but different `schemas` lists are different cache entries. the schema Register this once with `POST /v1/tenants/:t/schemas`. ``` namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "email" ty = "str" [[columns]] name = "tier" ty = "str" [[indexes]] name = "by_tier" columns = ["tier"] ``` seed data ``` # Seed customers across two tiers so the count below is non-trivial. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.customers/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "c_1", "email": "alice@example.com", "tier": "gold" }, { "id": "c_2", "email": "bob@example.com", "tier": "silver" }, { "id": "c_3", "email": "carol@example.com", "tier": "gold" }, { "id": "c_4", "email": "dan@example.com", "tier": "silver" } ]' ``` the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "how many gold-tier customers do I have?", "schemas": ["shop.customers"] }' ``` #### Python ``` result = db.ask( "how many gold-tier customers do I have?", schemas=["shop.customers"], ) print(result["rows"], result["cache"]) ``` #### TypeScript ``` const result = await db.ask( "how many gold-tier customers do I have?", { schemas: ["shop.customers"] } ); console.log(result.rows, result.cache); ``` #### Go ``` result, err := db.Ask(ctx, "how many gold-tier customers do I have?", "shop.customers", // variadic schemas ) if err != nil { /* handle */ } fmt.Println(result.Rows, result.Cache) ``` what you get back ``` { "rows": [ { "count": 2 } ], "cache": "miss" } ``` The shape is the same as the unhinted call - `rows` and `cache`. The difference is that the compiler walked one schema, not the whole tenant. how it works - The `schemas` list is part of the cache key - same question, different schemas list, different entry. - The compiler sees only the listed schemas in its catalog. Tables in other namespaces are invisible for this call. - Recognising "gold-tier" hits the `tier` column. Because there's a `by_tier` index, the plan becomes `Aggregate(count) → IndexScan(tier = "gold")`. - Fewer schemas = a smaller compile prompt and a smaller search space for join planning, so first-call latency drops. common mistakes - Listing only one side of a join. If the question implies a join across two schemas, both must be in the `schemas` list. Otherwise the compiler can't see the relation and falls back to `no_plan_compiled`. - Bare table name without namespace. Schemas are addressed as `namespace.table`. `"customers"` alone returns `400 unknown_schema`. - Expecting it to change rows. The hint changes which plans the compiler considers, not the results. If the hint is satisfiable, you get the same answer as the unhinted call - just compiled faster. --- # Return the compiled Plan tree - OriginChainDB Ask example Canonical source: https://originchaindb.com/docs/examples/ask/show-plan Sitemap last modified: 2026-09-13T20:18:50.000Z examples · ask · 3 / 6 # 3. Return the compiled Plan tree [← Ask examples](https://originchaindb.com/docs/examples/ask) what this does Adding `"show_plan": true` to the Ask body puts an extra `plan` field in the response: the executor plan tree the compiler produced - the same shape you'd pass to `/v1/query`. `show_plan` is a body field, not a query-string param. The engine reads it from the JSON body. A `?show_plan=true` on the URL is silently ignored. when to use it - You're about to put a question in a hot path and need to confirm it lands on the right index. - The same question is returning slower than you expect and you want to see what the planner picked. - You want to lift the plan and call `/v1/query` directly to skip the compile cost on every run. Leave it off in production application code - the plan tree is a few KB and you don't need it on every call. the schema ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer_id" ty = "str" [[columns]] name = "amount" ty = "f64" [[columns]] name = "status" ty = "str" [[indexes]] name = "by_status" columns = ["status"] ``` seed data ``` # Seed a handful of paid and pending orders so the plan has work to do. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "o_1", "customer_id": "c_1", "amount": 19.00, "status": "paid" }, { "id": "o_2", "customer_id": "c_2", "amount": 42.00, "status": "paid" }, { "id": "o_3", "customer_id": "c_1", "amount": 11.50, "status": "pending" } ]' ``` the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "count of paid orders", "schemas": ["shop.orders"], "show_plan": true }' ``` #### Python ``` result = db.ask( "count of paid orders", schemas=["shop.orders"], show_plan=True, ) print(result["rows"]) print(result["plan"]) # the executor plan tree ``` #### TypeScript ``` const result = await db.ask( "count of paid orders", { schemas: ["shop.orders"], show_plan: true } ); console.log(result.rows); console.log(result.plan); // the executor plan tree ``` #### Go ``` // Go SDK does not yet expose show_plan - call /ask via raw HTTP for the // plan tree, or use the Python / TypeScript SDK during plan inspection. result, err := db.Ask(ctx, "count of paid orders", "shop.orders") if err != nil { /* handle */ } fmt.Println(result.Rows) ``` what you get back ``` { "rows": [ { "count": 2 } ], "cache": "miss", "plan": { "Aggregate": { "group": [], "aggs": [ { "Count": { "alias": "count" } } ], "child": { "Filter": { "predicate": { "Eq": { "path": "status", "value": "paid" } }, "child": { "IndexScan": { "schema": "shop.orders", "index": "by_status" } } } } } } } ``` The `plan` field is the same JSON the engine's `/v1/query` endpoint accepts as input. Lift it, store it, replay it - the planner is deterministic for a fixed schema set. how it works - The compiler runs as usual. When `show_plan` is true, the resulting plan tree is serialised back into the response alongside the rows. - The plan here is `Aggregate(count) → Filter(status = paid) → IndexScan(by_status)`. The presence of `IndexScan` instead of plain `Scan` tells you the planner picked up the `by_status` index. - If you see `Scan` where you expected `IndexScan`, the index probably doesn't cover the predicate - add an index or rephrase. common mistakes - Putting `show_plan` in the URL. It's a body field. The query string is ignored. - Logging the plan tree on every request. It's a few KB and changes the response shape. Use it during development; turn it off in app code. - Assuming the plan is stable across schema changes. Adding or dropping an index re-plans. The previous plan tree is no longer the one the engine will run - re-fetch it. --- # Simple count question - OriginChainDB Ask example Canonical source: https://originchaindb.com/docs/examples/ask/simple-count Sitemap last modified: 2026-09-13T20:18:50.000Z examples · ask · 1 / 6 # 1. Simple count question [← Ask examples](https://originchaindb.com/docs/examples/ask) what this does POST /v1/tenants/:t/ask with a plain-English question in the `"nl"` field returns rows: the engine compiles the question into a Plan tree, runs it, and answers. This is the smallest Ask call - no schemas hint, no flags. when to use it - You're prototyping and don't yet know which schemas a question will touch. - The question is unambiguous against your registered schemas - one table is an obvious fit. - You want a quick sanity check that the Ask endpoint is wired up before you layer on schemas hints or show_plan. For production hot paths, pass a `schemas` list (next example) so the compiler doesn't have to consider every schema in the tenant. the schema Register this once with `POST /v1/tenants/:t/schemas` (Content-Type: `text/plain`). ``` namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "email" ty = "str" [[columns]] name = "region" ty = "str" ``` seed data Load three customers so the count returns a real number. ``` # One-time seed of three customers - skip if you've already loaded data. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.customers/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "c_1", "email": "alice@example.com", "region": "IN" }, { "id": "c_2", "email": "bob@example.com", "region": "US" }, { "id": "c_3", "email": "carol@example.com", "region": "DE" } ]' ``` the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "how many customers do I have?" }' ``` #### Python ``` result = db.ask("how many customers do I have?") print(result["rows"]) print(result["cache"]) # "miss" on first call ``` #### TypeScript ``` const result = await db.ask("how many customers do I have?"); console.log(result.rows); console.log(result.cache); // "miss" on first call ``` #### Go ``` result, err := db.Ask(ctx, "how many customers do I have?") if err != nil { /* handle */ } fmt.Println(result.Rows) fmt.Println(result.Cache) ``` what you get back ``` { "rows": [ { "count": 3 } ], "cache": "miss" } ``` Every Ask response has the same shape: `rows` (an array of objects keyed by column), `cache` (`"hit"`, `"miss"`, or `"skip"`), and an optional `plan` field when you asked for it. how it works - The Ask endpoint canonicalises the question (lowercase, collapse whitespace) and looks it up in the plan cache keyed by question + schemas list. - First call is a miss. The compiler walks the registered schemas, recognises "how many" as a COUNT, and produces a Plan tree like `Aggregate(count) → Scan(shop.customers)`. - The executor runs the plan and the count rolls up to one row. - The plan is stored under that cache key for next time - the actual rows are not cached, only the plan. common mistakes - Sending `"question"` instead of `"nl"`. The engine reads the body field `nl`. Anything else returns `400 missing_field`. - Expecting the rows array to be cached. Cache hits replay the compiled plan, not the result set. If the underlying rows changed since the previous call, you'll see the new count on a `"hit"`. - Ambiguous questions on a busy tenant. "How many customers" works if customers is the only obvious table. Once you have a `crm.customers` and a `shop.customers`, narrow with a schemas hint (next example). --- # Atomic multi-shape writes - OriginChainDB examples Canonical source: https://originchaindb.com/docs/examples/atomic-multi-shape Sitemap last modified: 2026-09-13T20:18:50.000Z examples · atomic # Atomic multi-shape writes [← All examples](https://originchaindb.com/docs/examples) There is no single "write everything" endpoint. Each shape has its own call - `/rows`, `/vector`, `/fts` - and each call is atomic by itself. The SDKs auto-attach an `Idempotency-Key` on every mutating call, so if one of the three fails you can retry just that one without re-writing the others. 5 recipes for writing the same logical object across multiple shapes - row, vector, full-text index, graph edge. Each recipe is on its own page with the schema TOML, the separate API calls, side-by-side cURL / Python / TypeScript / Go, and the common mistakes. [1 Product catalog (row + vector + FTS + graph edge) One product written as a row, a similarity vector, a BM25 full-text index, and a supplier edge declared by the schema. →](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog)[2 Knowledge-base article (row + vector + FTS) RAG and help-center search. Row plus an embedding of the body plus a keyword index on the body. No graph. →](https://originchaindb.com/docs/examples/atomic-multi-shape/kb-article)[3 Social user (row + vector + graph self-relation) Profile row, profile embedding for similar-user lookup, and a follows self-relation that writes both directions. →](https://originchaindb.com/docs/examples/atomic-multi-shape/social-user)[4 Audit event (row + FTS) Simplest atomic write. Append-only event row plus a free-text index over the payload for compliance search. →](https://originchaindb.com/docs/examples/atomic-multi-shape/audit-event)[5 Agent memory (row + vector + FTS + metadata filter) Long-term memory for agents. Row, embedding with agent_id + session_id metadata, and a keyword index over the text. →](https://originchaindb.com/docs/examples/atomic-multi-shape/agent-memory) --- # Agent memory: row, vector, FTS, metadata - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/atomic-multi-shape/agent-memory Sitemap last modified: 2026-09-13T20:18:50.000Z examples · atomic · 5 / 5 # 5. Agent memory (row + vector + FTS + metadata filter) [← Atomic multi-shape](https://originchaindb.com/docs/examples/atomic-multi-shape) what this does Save one piece of agent memory three ways - the structured row in `agent.memories`, an embedding for "what did this agent learn that's similar to the current query", and a keyword index over the text. The vector put carries `agent_id` and `session_id` as metadata, so similarity search can be filtered to a single agent and session at retrieval time. when to use it - Chat assistants with long-term memory: before each turn, retrieve the top-k similar past memories for the current agent and session. - Autonomous agents that need to recall prior tool calls, plans, or facts. - Any system where the same text store is queried by both vector similarity (for recall) and keywords (for exact-phrase grep). the schema The composite index covers the most common read shape - "newest memories for this agent in this session". ``` # agent/memories.toml namespace = "agent" table = "memories" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "agent_id" ty = "str" required = true [[columns]] name = "session_id" ty = "str" required = true [[columns]] name = "text" ty = "str" required = true [[columns]] name = "ts_ms" ty = "u64" required = true # Composite index supports "all memories for this agent in this session, # newest first" without a full scan. [[indexes]] name = "by_agent_session_time" columns = ["agent_id", "session_id", "ts_ms"] ``` call 1 of 3 - the memory row #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/agent.memories" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "mem-77f2-001", "agent_id": "agt-sales-east", "session_id": "thr-77f2", "text": "Customer prefers contact between 09:00 and 11:00 IST.", "ts_ms": 1749500120000 }' ``` #### Python ``` db.rows.put("agent.memories", { "id": "mem-77f2-001", "agent_id": "agt-sales-east", "session_id": "thr-77f2", "text": "Customer prefers contact between 09:00 and 11:00 IST.", "ts_ms": 1749500120000, }) ``` #### TypeScript ``` // The TypeScript SDK does not wrap row writes yet // (shipping in the next release). Use `fetch` for now. await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/agent.memories`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "mem-77f2-001", agent_id: "agt-sales-east", session_id: "thr-77f2", text: "Customer prefers contact between 09:00 and 11:00 IST.", ts_ms: 1749500120000, }), }); ``` #### Go ``` // The Go SDK does not wrap row writes yet // (shipping in the next release). Use net/http for now. body, _ := json.Marshal(map[string]any{ "id": "mem-77f2-001", "agent_id": "agt-sales-east", "session_id": "thr-77f2", "text": "Customer prefers contact between 09:00 and 11:00 IST.", "ts_ms": uint64(1749500120000), }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/agent.memories", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` call 2 of 3 - the embedding with metadata `agent_id` and `session_id` go in `metadata`, not just in the row. Without them on the vector, you can't restrict `/vector/topk` to a single agent at search time, and one agent's memories leak into another's recall. #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/vector/agent.memories/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "mem-77f2-001", "embedding": [0.0144, -0.0681, 0.0398, /* ... 768 floats ... */], "dim": 768, "metric": "cosine", "metadata": { "agent_id": "agt-sales-east", "session_id": "thr-77f2" } }' ``` #### Python ``` # Metadata enables "find similar memories for THIS agent in THIS session" # at search time, without scanning every other agent's memories. db.vector.put( "agent.memories", "mem-77f2-001", embedding_768d, metadata={ "agent_id": "agt-sales-east", "session_id": "thr-77f2", }, ) ``` #### TypeScript ``` await db.vectorPut("agent.memories", { id: "mem-77f2-001", embedding: embedding768d, dim: 768, metric: "cosine", metadata: { agent_id: "agt-sales-east", session_id: "thr-77f2", }, }); ``` #### Go ``` err := db.VectorPut(ctx, "agent.memories", originchain.VectorPutRequest{ ID: "mem-77f2-001", Embedding: embedding768d, Dim: 768, Metric: "cosine", Metadata: map[string]any{ "agent_id": "agt-sales-east", "session_id": "thr-77f2", }, }) ``` call 3 of 3 - the keyword index #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/fts/agent.memories/text" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "mem-77f2-001", "text": "Customer prefers contact between 09:00 and 11:00 IST." }' ``` #### Python ``` db.fts.index( "agent.memories", "text", doc_id="mem-77f2-001", text="Customer prefers contact between 09:00 and 11:00 IST.", ) ``` #### TypeScript ``` await db.ftsIndex("agent.memories", "text", { doc_id: "mem-77f2-001", text: "Customer prefers contact between 09:00 and 11:00 IST.", }); ``` #### Go ``` err := db.FTSIndex(ctx, "agent.memories", "text", originchain.FTSIndexRequest{ DocID: "mem-77f2-001", Text: "Customer prefers contact between 09:00 and 11:00 IST.", }) ``` about atomicity The three calls are separate. There is no single "write everything" endpoint. Each call is atomic by itself. The SDKs auto-attach an `Idempotency-Key` on every mutating call, so if the vector put fails after the row succeeded, retry just that one - re-doing the row write would not duplicate the memory. common mistakes - Not partitioning by `agent_id`. If the vector put doesn't carry the agent metadata, every `topk` ranks across every agent's memories. One sales agent ends up "recalling" what a support agent learned about a different customer. - Storing huge memory blocks as one row. Memories should be sentence- or paragraph-sized. A whole conversation in one row gives mush at retrieval time. Split before writing. - Embedding without normalizing first. If you compare a fresh embedding (model A, length-normalized) to a stored one (model B, raw), cosine scores are meaningless. Pin one model + one normalization step at write and read time. - Never expiring old memories. Agent memory tables grow without bound. Build a sweeper that deletes by `ts_ms` below a threshold, and remember to `DELETE` the vector and FTS doc too - row deletion does not fan out. --- # Audit event - row + FTS - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/atomic-multi-shape/audit-event Sitemap last modified: 2026-09-13T20:18:50.000Z examples · atomic · 4 / 5 # 4. Audit event (row + FTS) [← Atomic multi-shape](https://originchaindb.com/docs/examples/atomic-multi-shape) what this does One event row in `audit.events` plus a BM25 full-text index over the human-readable payload - the simplest multi-shape recipe. No vector, no graph, just a structured event and searchable text. when to use it - Compliance / SOC 2 audit logs you'll search by free text later ("who changed the carbon_plate column?"). - Activity feeds where the structured fields (`actor_id`, `action`, `ts_ms`) are the filter and the payload text is what humans read. - Application logs where you want both fast time-range scans and grep-style search over the message. the schema Two indexes: one for "find all events by this actor in this window" and one for global time-range queries. ``` # audit/events.toml namespace = "audit" table = "events" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "actor_id" ty = "str" required = true [[columns]] name = "action" ty = "str" required = true [[columns]] name = "payload_text" ty = "str" [[columns]] name = "ts_ms" ty = "u64" required = true [[indexes]] name = "by_actor_and_time" columns = ["actor_id", "ts_ms"] [[indexes]] name = "by_time" columns = ["ts_ms"] ``` call 1 of 2 - the event row #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/audit.events" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "evt-2026-06-10-00042", "actor_id": "user.alice", "action": "schema.update", "payload_text": "Added column carbon_plate to shop.products", "ts_ms": 1749500000000 }' ``` #### Python ``` db.rows.put("audit.events", { "id": "evt-2026-06-10-00042", "actor_id": "user.alice", "action": "schema.update", "payload_text": "Added column carbon_plate to shop.products", "ts_ms": 1749500000000, }) ``` #### TypeScript ``` // The TypeScript SDK does not wrap row writes yet // (shipping in the next release). Use `fetch` for now. await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/audit.events`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "evt-2026-06-10-00042", actor_id: "user.alice", action: "schema.update", payload_text: "Added column carbon_plate to shop.products", ts_ms: 1749500000000, }), }); ``` #### Go ``` // The Go SDK does not wrap row writes yet // (shipping in the next release). Use net/http for now. body, _ := json.Marshal(map[string]any{ "id": "evt-2026-06-10-00042", "actor_id": "user.alice", "action": "schema.update", "payload_text": "Added column carbon_plate to shop.products", "ts_ms": uint64(1749500000000), }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/audit.events", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` call 2 of 2 - the keyword index Index the sanitized payload text. Same `doc_id` as the row's primary key so search hits join back to the event. #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/fts/audit.events/payload_text" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "evt-2026-06-10-00042", "text": "Added column carbon_plate to shop.products" }' ``` #### Python ``` # Sanitize first - never index raw PII. text = redact_pii(payload_text) # your redaction step db.fts.index( "audit.events", "payload_text", doc_id="evt-2026-06-10-00042", text=text, ) ``` #### TypeScript ``` // Sanitize first - never index raw PII. const text = redactPII(payloadText); await db.ftsIndex("audit.events", "payload_text", { doc_id: "evt-2026-06-10-00042", text: text, }); ``` #### Go ``` // Sanitize first - never index raw PII. text := RedactPII(payloadText) err := db.FTSIndex(ctx, "audit.events", "payload_text", originchain.FTSIndexRequest{ DocID: "evt-2026-06-10-00042", Text: text, }) ``` about atomicity The row write and the FTS index are separate calls. There is no single "write everything" endpoint. Each is atomic by itself. The SDKs auto-attach an `Idempotency-Key`, so if the FTS index fails after the row succeeded, you can safely retry just the FTS call - re-doing the row write would not duplicate the event. common mistakes - Indexing PII without redaction. Anything that lands in the FTS index is searchable forever. Sanitize emails, phone numbers, tokens, internal IDs before calling `db.fts.index` - never after. - Mutating audit events. Audit logs should be append-only. Don't re-put the row to "correct" an event - emit a new event that supersedes it. - Embedding raw JSON in `payload_text`. FTS tokenization treats `` and `"` as noise. Render the payload as a human sentence first so the index actually has searchable words. - Forgetting an `id`. If the row write generates an id server-side and the FTS call uses a different one, search hits won't join back. Set the id client-side (UUID or a sortable like ULID) and use the same string for both calls. --- # Knowledge-base article - row + vector + FTS - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/atomic-multi-shape/kb-article Sitemap last modified: 2026-09-13T20:18:50.000Z examples · atomic · 2 / 5 # 2. Knowledge-base article (row + vector + FTS) [← Atomic multi-shape](https://originchaindb.com/docs/examples/atomic-multi-shape) what this does Save one help-center article as three shapes - the structured row in `kb.articles`, an embedding of the title and body for semantic similarity, and a BM25 keyword index on the body. No graph edges - this is the simplest recipe that still covers a real RAG / search backend. when to use it - Help-center search backends. Keyword search handles "exact phrase" queries; vector search handles "what does this user mean". - RAG retrieval - rank candidates by vector similarity, optionally re-rank with FTS, then fetch the body from the row store. - Any corpus where you want both lexical and semantic recall over the same documents. the schema Plain row schema - no `[[relations]]` because there's no graph edge to write. ``` # kb/articles.toml namespace = "kb" table = "articles" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "title" ty = "str" required = true [[columns]] name = "body" ty = "str" required = true [[columns]] name = "url" ty = "str" [[columns]] name = "created_ms" ty = "u64" [[indexes]] name = "by_created" columns = ["created_ms"] ``` call 1 of 3 - the row #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/kb.articles" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "kb-2026-001", "title": "How atomic multi-shape writes work", "body": "Each shape (row, vector, FTS, graph) has its own endpoint. Every call is atomic individually. Idempotency keys make retries safe.", "url": "/docs/concepts/atomic-multi-shape", "created_ms": 1747900000000 }' ``` #### Python ``` db.rows.put("kb.articles", { "id": "kb-2026-001", "title": "How atomic multi-shape writes work", "body": "Each shape (row, vector, FTS, graph) has its own endpoint. Every call is atomic individually. Idempotency keys make retries safe.", "url": "/docs/concepts/atomic-multi-shape", "created_ms": 1747900000000, }) ``` #### TypeScript ``` // The TypeScript SDK does not wrap row writes yet // (shipping in the next release). Use `fetch` for now. await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/kb.articles`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "kb-2026-001", title: "How atomic multi-shape writes work", body: "Each shape (row, vector, FTS, graph) has its own endpoint. Every call is atomic individually. Idempotency keys make retries safe.", url: "/docs/concepts/atomic-multi-shape", created_ms: 1747900000000, }), }); ``` #### Go ``` // The Go SDK does not wrap row writes yet // (shipping in the next release). Use net/http for now. body, _ := json.Marshal(map[string]any{ "id": "kb-2026-001", "title": "How atomic multi-shape writes work", "body": "Each shape (row, vector, FTS, graph) has its own endpoint. Every call is atomic individually. Idempotency keys make retries safe.", "url": "/docs/concepts/atomic-multi-shape", "created_ms": uint64(1747900000000), }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/kb.articles", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` call 2 of 3 - the embedding (title + body concatenated) Embed both fields together. A title-only embedding misses everything the body says, and that's where most of the meaningful tokens live. #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/vector/kb.articles/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "kb-2026-001", "embedding": [0.0211, -0.0612, 0.0341, /* ... 768 floats ... */], "dim": 768, "metric": "cosine" }' ``` #### Python ``` # Embed the title and body together so semantic search hits both. text = f"{title}\n\n{body}" embedding_768d = embed(text) # any embedding model db.vector.put( "kb.articles", "kb-2026-001", embedding_768d, ) ``` #### TypeScript ``` // Embed title + body together. embedding768d is your number[] of length 768. await db.vectorPut("kb.articles", { id: "kb-2026-001", embedding: embedding768d, dim: 768, metric: "cosine", }); ``` #### Go ``` // Embed title + body together. embedding768d is your []float32 of length 768. err := db.VectorPut(ctx, "kb.articles", originchain.VectorPutRequest{ ID: "kb-2026-001", Embedding: embedding768d, Dim: 768, Metric: "cosine", }) ``` call 3 of 3 - the keyword index (body only) Index the body for BM25 keyword search. Help-center queries are usually keyword-shaped ("install on Windows"). #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/fts/kb.articles/body" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "kb-2026-001", "text": "Each shape (row, vector, FTS, graph) has its own endpoint. Every call is atomic individually. Idempotency keys make retries safe." }' ``` #### Python ``` db.fts.index( "kb.articles", "body", doc_id="kb-2026-001", text=body, ) ``` #### TypeScript ``` await db.ftsIndex("kb.articles", "body", { doc_id: "kb-2026-001", text: body, }); ``` #### Go ``` err := db.FTSIndex(ctx, "kb.articles", "body", originchain.FTSIndexRequest{ DocID: "kb-2026-001", Text: body, }) ``` about atomicity The three calls are separate. There is no single "write everything" endpoint. Each call is atomic by itself. The SDKs auto-attach an `Idempotency-Key` on every mutating call, so if the FTS call fails after the row and vector succeeded, retry just the FTS one - re-doing the row write would not duplicate it. common mistakes - Embedding only the title. Concatenate title + body before embedding so semantic similarity fires on body content, not just the headline. - Indexing the title in FTS but not the body. The opposite mistake. The body is where the searchable keywords live. - Forgetting to re-index on update. If you edit an article, you have to re-put the row, re-put the vector, and re-index the FTS field. None of the three rides along with the others. - Embedding huge bodies as one vector. Past ~1k tokens, semantic similarity gets muddy. For long articles, chunk the body, write one row per chunk with a parent `article_id`, and embed each chunk separately. --- # Product catalog: row, vector, FTS, graph - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog Sitemap last modified: 2026-09-13T20:18:50.000Z examples · atomic · 1 / 5 # 1. Product catalog (row + vector + FTS + graph edge) [← Atomic multi-shape](https://originchaindb.com/docs/examples/atomic-multi-shape) what this does Save one product across four shapes - the structured row in `shop.products`, an embedding of the name + description for similarity search, a BM25 full-text index on the description for keyword search, and a graph edge to its supplier. The supplier edge is the only one you don't write by hand: declaring `supplier_id` as a `[[relations]]` column in the schema makes the row write also create the edge in the same commit. when to use it - E-commerce catalogs where one product needs to be findable by keyword, by similarity, and by a graph walk (e.g. "all products from suppliers in this region"). - Any domain object you want to query in more than one shape - the three calls run once per object and then every read path works. the schema Push this once with `/v1/tenants//schemas` before the first write. Note the `[[relations]]` block - that's what makes the supplier edge automatic. ``` # shop/products.toml namespace = "shop" table = "products" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" required = true [[columns]] name = "description" ty = "str" [[columns]] name = "price_cents" ty = "i64" [[columns]] name = "category" ty = "str" [[columns]] name = "supplier_id" ty = "str" required = true [[indexes]] name = "by_category" columns = ["category"] # Declaring supplier_id as a relation makes the row write also # create a graph edge from this product to its supplier. [[relations]] name = "supplied_by" from_col = "supplier_id" bidirectional = true [relations.target] namespace = "shop" table = "suppliers" pk = "id" ``` call 1 of 3 - the row (also writes the supplier edge) #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/shop.products" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "sku-9281", "name": "Trailblazer 2", "description": "All-terrain running shoe with carbon-plate forefoot.", "price_cents": 12900, "category": "running-shoes", "supplier_id": "sup-acme-shoes" }' ``` #### Python ``` db.rows.put("shop.products", { "id": "sku-9281", "name": "Trailblazer 2", "description": "All-terrain running shoe with carbon-plate forefoot.", "price_cents": 12900, "category": "running-shoes", "supplier_id": "sup-acme-shoes", }) # The supplied_by edge to sup-acme-shoes is written in the same commit - # the [[relations]] block in the schema makes that automatic. ``` #### TypeScript ``` // The TypeScript SDK does not wrap row writes yet // (shipping in the next release). Use `fetch` for now - // it is exactly the same HTTP call the SDK will make. await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/shop.products`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "sku-9281", name: "Trailblazer 2", description: "All-terrain running shoe with carbon-plate forefoot.", price_cents: 12900, category: "running-shoes", supplier_id: "sup-acme-shoes", }), }); ``` #### Go ``` // The Go SDK does not wrap row writes yet // (shipping in the next release). Use net/http for now. body, _ := json.Marshal(map[string]any{ "id": "sku-9281", "name": "Trailblazer 2", "description": "All-terrain running shoe with carbon-plate forefoot.", "price_cents": 12900, "category": "running-shoes", "supplier_id": "sup-acme-shoes", }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/shop.products", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` call 2 of 3 - the embedding The `id` here is the same primary key as the row. Same key = same logical product across shapes. #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/vector/shop.products/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "sku-9281", "embedding": [0.0124, -0.0883, 0.0451, /* ... 768 floats ... */], "dim": 768, "metric": "cosine", "metadata": { "category": "running-shoes" } }' ``` #### Python ``` db.vector.put( "shop.products", "sku-9281", embedding_768d, # name+description embedded together metadata={ "category": "running-shoes" }, ) ``` #### TypeScript ``` await db.vectorPut("shop.products", { id: "sku-9281", embedding: embedding768d, dim: 768, metric: "cosine", metadata: { category: "running-shoes" }, }); ``` #### Go ``` err := db.VectorPut(ctx, "shop.products", originchain.VectorPutRequest{ ID: "sku-9281", Embedding: embedding768d, Dim: 768, Metric: "cosine", Metadata: map[string]any{ "category": "running-shoes", }, }) ``` call 3 of 3 - the full-text index `doc_id` matches the row's primary key so search results can be joined back to the row. #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "sku-9281", "text": "Trailblazer 2 all-terrain running shoe with carbon-plate forefoot" }' ``` #### Python ``` db.fts.index( "shop.products", "description", doc_id="sku-9281", text="Trailblazer 2 all-terrain running shoe with carbon-plate forefoot", ) ``` #### TypeScript ``` await db.ftsIndex("shop.products", "description", { doc_id: "sku-9281", text: "Trailblazer 2 all-terrain running shoe with carbon-plate forefoot", }); ``` #### Go ``` err := db.FTSIndex(ctx, "shop.products", "description", originchain.FTSIndexRequest{ DocID: "sku-9281", Text: "Trailblazer 2 all-terrain running shoe with carbon-plate forefoot", }) ``` about atomicity The three calls are separate. There is no single "write everything" endpoint. Each call is atomic by itself - the row write either commits or doesn't, the vector put either commits or doesn't. Every mutating call gets an auto-attached `Idempotency-Key` from the SDK, so if the FTS call fails after the row and vector succeeded, you can safely retry just the FTS one without re-writing the others. common mistakes - Forgetting one of the three calls. If you skip the vector put, similarity search won't return this product. If you skip the FTS index, keyword search won't. The row write doesn't fan out to the other shapes. - Updating one without re-doing the others. If you re-put the row with a new description, the vector and FTS index still point at the old text until you re-put / re-index them too. - Mismatched IDs. The `id` in the row write, the vector put, and the FTS `doc_id` must all be the same string. If they drift, you can't join search hits back to the row. - Writing the supplier edge by hand. Don't. Let the `[[relations]]` block do it. There is no edge-write endpoint to call instead — edges are derived from the row write, so a second "write the edge" step is not something you can do, and not something you need to do. --- # Social user: row, vector, follow graph - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/atomic-multi-shape/social-user Sitemap last modified: 2026-09-13T20:18:50.000Z examples · atomic · 3 / 5 # 3. Social user (row + vector + graph self-relation) [← Atomic multi-shape](https://originchaindb.com/docs/examples/atomic-multi-shape) what this does Save a user as a row in `social.users`, embed their bio for "find similar users", and model the follow graph as a `follows` list column on that same row. A `[[relations]]` self-relation turns each element of the list into a graph edge (`alice → bob`, `alice → carol`), and `bidirectional = true` maintains the reverse side so "who follows me?" is a graph read. when to use it - "People you may know" - similar-user lookup over bio embeddings, intersected with friends-of-friends from the follow graph. - Reverse-adjacency reads like "who follows me?" are a graph walk, not an index scan, because the bidirectional relation maintains both sides. - Any social, professional, or collaboration network with self-relations between users. the schema (one table) Push this once. The `follows` list column plus the self-relation is the whole graph - no separate edge table to keep in sync. ``` # social/users.toml - one table models the profile AND the graph. namespace = "social" table = "users" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "handle" ty = "str" required = true [[columns]] name = "bio" ty = "str" [[columns]] name = "joined_ms" ty = "u64" # The follow graph lives in a list column on the row. Each element is a # user id this person follows. [[columns]] name = "follows" ty = "list" element = "str" [[indexes]] name = "by_handle" columns = ["handle"] # A self-relation turns the 'follows' list into graph edges: one forward # edge per element (alice -> bob, alice -> carol). bidirectional = true also # maintains the reverse side, so "who follows bob?" is a graph read, not a scan. [[relations]] name = "following" from_col = "follows" bidirectional = true [relations.target] namespace = "social" table = "users" pk = "id" ``` call 1 of 3 - the user row (profile + follows) The `follows` array lands as graph edges in the same write. To follow or unfollow later, re-put the row with the updated list - the edges are recomputed from the new list on each write. #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/social.users" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "alice", "handle": "@alice", "bio": "Mechanical engineer. Trail runner. Reading sci-fi.", "joined_ms": 1746180851000, "follows": ["bob", "carol"] }' ``` #### Python ``` db.rows.put("social.users", { "id": "alice", "handle": "@alice", "bio": "Mechanical engineer. Trail runner. Reading sci-fi.", "joined_ms": 1746180851000, "follows": ["bob", "carol"], }) ``` #### TypeScript ``` // The TypeScript SDK writes rows via SQL INSERT (typed row helpers // ship in a later release). Use `fetch` for the row-shaped write: await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/social.users`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "alice", handle: "@alice", bio: "Mechanical engineer. Trail runner. Reading sci-fi.", joined_ms: 1746180851000, follows: ["bob", "carol"], }), }); ``` #### Go ``` // The Go SDK writes rows via SQL INSERT (typed row helpers ship // in a later release). Use net/http for the row-shaped write: body, _ := json.Marshal(map[string]any{ "id": "alice", "handle": "@alice", "bio": "Mechanical engineer. Trail runner. Reading sci-fi.", "joined_ms": uint64(1746180851000), "follows": []string{"bob", "carol"}, }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/social.users", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` call 2 of 3 - the profile embedding Use the same `id` as the user row. Now `/vector/social.users/topk` returns similar users by bio. #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/vector/social.users/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "alice", "embedding": [0.0182, -0.0712, 0.0419, /* ... 768 floats ... */], "dim": 768, "metric": "cosine" }' ``` #### Python ``` # Embed the bio (and optionally handle) so "find similar users" works. db.vector.put( "social.users", "alice", embedding_768d, ) ``` #### TypeScript ``` await db.vectorPut("social.users", { id: "alice", embedding: embedding768d, dim: 768, metric: "cosine", }); ``` #### Go ``` err := db.VectorPut(ctx, "social.users", originchain.VectorPutRequest{ ID: "alice", Embedding: embedding768d, Dim: 768, Metric: "cosine", }) ``` call 3 of 3 - traverse the follow graph `neighbors` walks one hop forward along the relation (who alice follows); `reverse` walks the back-edge (who follows bob). Both are direct adjacency-list lookups - no scan, no join. #### cURL ``` # Who does alice follow? (one hop forward along 'following') curl "$ORIGINCHAIN_URL/v1/tenants/$T/graph/social.users/neighbors?rel=following&pk=alice" \ -H "Authorization: Bearer $OC_TOKEN" # → ["bob", "carol"] # Who follows bob? (reverse side, maintained by bidirectional = true) curl "$ORIGINCHAIN_URL/v1/tenants/$T/graph/social.users/reverse?rel=following&pk=bob" \ -H "Authorization: Bearer $OC_TOKEN" # → ["alice"] ``` #### Python ``` # Who does alice follow? db.graph.neighbors("social.users", rel="following", pk="alice") # -> ["bob", "carol"] # Who follows bob? db.graph.reverse_neighbors("social.users", rel="following", pk="bob") # -> ["alice"] ``` #### TypeScript ``` // Who does alice follow? await oc.graph.neighbors("social.users", { rel: "following", pk: "alice" }); // ["bob","carol"] // Who follows bob? await oc.graph.reverseNeighbors("social.users", { rel: "following", pk: "bob" }); // ["alice"] ``` #### Go ``` // Who does alice follow? db.Graph().Neighbors(ctx, "social.users", originchain.NeighborsRequest{Rel: "following", PK: "alice"}) // ["bob","carol"] // Who follows bob? db.Graph().ReverseNeighbors(ctx, "social.users", originchain.NeighborsRequest{Rel: "following", PK: "bob"}) // ["alice"] ``` about atomicity The user row and the profile embedding are separate calls - each atomic by itself. There is no single "write everything" endpoint. The row write (including its `follows` edges) commits atomically, and the SDKs auto-attach an `Idempotency-Key` on every mutating call, so if the vector put fails after the row succeeded, retry just that one. common mistakes - Modeling follows as a composite-PK edge table. A `follows` table keyed `[follower_id, followee_id]` can't be traversed with `neighbors(pk="alice")` - the source of each edge is the whole composite key, not a single user. Put the relation on the `users` row via a `follows` list instead. - Forgetting the reverse direction. Set `bidirectional = true` so "who follows bob?" is a `reverse` read. Without it you can only walk forward. - Partial re-put on follow/unfollow. Edges are recomputed from the `follows` list on each write, so re-put the row with the full new list, not just the delta. - Updating handle without re-embedding. If you embed the handle and let users edit it, the vector goes stale until you re-put. Either re-put on every profile edit or only embed the bio. --- # Elasticsearch @elastic client examples - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/elasticsearch Sitemap last modified: 2026-09-13T20:18:50.000Z examples · elasticsearch # Elasticsearch examples [← All examples](https://originchaindb.com/docs/examples) Copy-paste recipes for the Elasticsearch API, using the official @elastic client. Each one reads from or writes to /v1/tenants/:t/es/. New here? Set the client up in [Connect an Elasticsearch client](https://originchaindb.com/docs/connect-elasticsearch), or see the raw endpoints in the [HTTP API reference](https://originchaindb.com/docs/api#elasticsearch). These use the Node client; the same requests work from any Elasticsearch 7.x client or straight over HTTP. Writes commit with the row, so a document is searchable immediately, and row-level security and column masking apply to every query. ## 1Create an index with a mapping Declare the fields up front. Types are the Elasticsearch types your client already sends. ``` await es.indices.create({ index: 'shop.products', mappings: { properties: { name: { type: 'text' }, brand: { type: 'keyword' }, price: { type: 'integer' } } } }) ``` ## 2Index a document Create or replace a document by _id. It is searchable the instant the call returns — there is no refresh interval to wait on. ``` await es.index({ index: 'shop.products', id: 'sku-8842', document: { name: 'Carbon Marathon', brand: 'Aero', price: 149 } }) ``` ## 3Bulk-load many documents The fast path for ingest and backfills. One NDJSON action line per document (the client builds it for you from the operations array). ``` await es.bulk({ operations: [ { index: { _index: 'shop.products', _id: 'sku-1207' } }, { name: 'Trail 24', brand: 'Aero', price: 89 }, { index: { _index: 'shop.products', _id: 'sku-3355' } }, { name: 'City Runner', brand: 'Metro', price: 72 } ]}) ``` ## 4Partial update (merge) Send only the fields that change. Everything else on the document is preserved. ``` await es.update({ index: 'shop.products', id: 'sku-8842', doc: { price: 139 } // name, brand, ... untouched }) ``` ## 5Full-text match Run the Query DSL you already write. match analyzes the query the same way the field was indexed. ``` await es.search({ index: 'shop.products', query: { match: { name: 'marathon' } } }) ``` ## 6Bool query with a filter Combine a scoring must with non-scoring filter clauses — a term and a numeric range. ``` await es.search({ index: 'shop.products', query: { bool: { must: [{ match: { name: 'runner' } }], filter: [{ term: { brand: 'Metro' } }, { range: { price: { lte: 100 } } }] } } }) ``` ## 7Aggregate (terms) Set size: 0 to skip the hits and get just the buckets — the shape Kibana and Grafana panels use. ``` await es.search({ index: 'shop.products', size: 0, aggs: { by_brand: { terms: { field: 'brand' } } } }) ``` ## 8Count matches A count without the documents. ``` await es.count({ index: 'shop.products', query: { term: { brand: 'Aero' } } }) ``` ## 9Paginate with search_after Deep paging within the 10,000-hit window. Sort by a field plus _id, then pass the previous page’s last sort tuple. ``` const page1 = await es.search({ index: 'shop.products', size: 20, sort: [{ price: 'asc' }, { _id: 'asc' }], query: { match_all: {} } }) const last = page1.hits.hits.at(-1).sort // e.g. [72, 'sku-3355'] await es.search({ index: 'shop.products', size: 20, sort: [{ price: 'asc' }, { _id: 'asc' }], search_after: last, query: { match_all: {} } }) ``` ## 10Delete by query (bounded) max_docs caps how many are deleted. An unbounded whole-table delete is refused rather than run — use _reindex into a fresh index for that. ``` await es.deleteByQuery({ index: 'shop.products', max_docs: 100, query: { term: { brand: 'Aero' } } }) ``` ## 11Read a document back Fetch a document by id. _version is the number the write reported, so you can carry it straight into a conditional write. _source filtering works here too, and _source/:id returns the bare document with no envelope. ``` await es.get({ index: 'shop.products', id: 'sku-8842' }) // -> { _index, _id, _version: 2, found: true, _source: { ... } } await es.getSource({ index: 'shop.products', id: 'sku-8842', _source_includes: 'price' }) // -> { price: 139 } await es.exists({ index: 'shop.products', id: 'sku-8842' }) // true / false ``` ## 12Bucket by calendar day Group documents into calendar buckets over a date field. The value can be an ISO-8601 string or epoch milliseconds — both bucket the same way. Buckets are UTC and calendar-aligned, so calendar_interval is what you pass; fixed_interval is refused rather than approximated with the wrong boundaries. ``` await es.search({ index: 'shop.orders', size: 0, aggs: { per_day: { date_histogram: { field: 'placed_at', calendar_interval: 'day', time_zone: 'UTC' } } } }) ``` ## 13Named filter buckets One bucket per named query, each counting the documents that match both your search and that filter, with sub-aggregations folded over exactly those documents. other_bucket_key adds a bucket for everything that matched none of them. Send a filters aggregation in its own request: it is answered with one search per filter, so it does not share a request with other top-level aggregations. ``` await es.search({ index: 'shop.products', size: 0, aggs: { by_price: { filters: { other_bucket_key: 'mid_range', filters: { clearance: { range: { price: { lt: 50 } } }, premium: { range: { price: { gte: 200 } } } } }, aggs: { avg_price: { avg: { field: 'price' } } } } } }) ``` ## 14Top documents inside an aggregation Return the best matching documents, with their _source, from inside an aggregation. It works on its own at the top level, or nested under a bucket aggregation so every bucket carries its own examples. ``` await es.search({ index: 'shop.products', size: 0, query: { match: { name: 'runner' } }, aggs: { best: { top_hits: { size: 3 } } } }) // nested: one top hit per brand aggs: { by_brand: { terms: { field: 'brand' }, aggs: { top: { top_hits: { size: 1 } } } } } ``` ## 15Did you mean (term suggester) Correct a typo against the words actually indexed in a field. suggest_mode defaults to missing, which corrects only words the index does not already hold; always corrects every word. Each option carries the corrected text and a score. The term suggester is the one available — phrase and completion suggesters are refused by name rather than answered approximately. ``` await es.search({ index: 'shop.products', size: 0, suggest: { did_you_mean: { text: 'marathn', term: { field: 'name' } } } }) // suggest.did_you_mean[0].options // -> [ { text: 'marathon', score: 0.875 } ] ``` ## 16Concurrency-safe bulk writes Carry if_seq_no and if_primary_term on a bulk action and that write lands only if nobody changed the document first. A stale precondition comes back as a per-item 409 and every other item in the batch still applies, so a losing writer retries one document rather than the whole load. ``` await es.bulk({ operations: [ { index: { _index: 'shop.products', _id: 'sku-8842', if_seq_no: 1, if_primary_term: 1 } }, { name: 'Carbon Marathon', brand: 'Aero', price: 129 } ]}) // stale precondition -> items[0].index.status === 409 // version_conflict_engine_exception ``` Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. OriginChainDB is not affiliated with, endorsed by, or sponsored by Elasticsearch B.V. OriginChainDB implements a compatible HTTP API so that existing Elasticsearch clients can talk to it; it does not distribute Elasticsearch software. --- # HTTP error codes and responses - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/errors Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors # Error responses [← All examples](https://originchaindb.com/docs/examples) 19 canonical errors, one per page. Every error response from the OriginChainDB API uses the same JSON shape: ``` { "error": "code", // short machine-readable string "message": "human", // plain-English explanation "retry": false // true if the same request might succeed later } ``` Some errors carry extra typed fields - `purchase_url` on a 402, `retry_after_seconds` on a 429, `current_version` on an OCC 409, `trace_id` on a 500. The individual pages show those. Every request example on them reads three environment variables: `OC_HOST`, `OC_TENANT` and `OC_TOKEN`. ## 400 - bad request The request reached the server, parsed, and was rejected before the engine ran it. Fix the request and resend. [1 400 - malformed JSON body The request body failed to parse as JSON. Trailing comma, unterminated string, wrong content-type. →](https://originchaindb.com/docs/examples/errors/400-malformed-json)[2 400 - unknown table (schema_not_found via SQL) The SQL referenced a table that is not declared in any schema for this tenant. →](https://originchaindb.com/docs/examples/errors/400-unknown-table)[3 400 - unknown column in SELECT/WHERE The translator could not resolve a column reference against the table's manifest. →](https://originchaindb.com/docs/examples/errors/400-unknown-column)[4 400 - column type mismatch on write A value did not match the column's declared type. String into a float column, etc. →](https://originchaindb.com/docs/examples/errors/400-type-mismatch)[5 400 - invalid vector mode (only fast / high_recall) The topk mode field has only two legal values. Anything else returns 400. →](https://originchaindb.com/docs/examples/errors/400-invalid-vector-mode) ## 401 / 402 / 403 - auth, billing, permission The request was authenticated or authorised incorrectly, or it asks for a paid feature the tenant hasn't enabled. [6 401 - missing Authorization header No Authorization header was sent at all. The router rejects before reaching the handler. →](https://originchaindb.com/docs/examples/errors/401-missing-bearer)[7 401 - malformed bearer token The header was present but did not parse as 'Bearer '. →](https://originchaindb.com/docs/examples/errors/401-malformed-bearer)[8 401 - token revoked or expired The token parsed but has been revoked from the admin console or is past its expiry. →](https://originchaindb.com/docs/examples/errors/401-expired)[9 402 - addon required (paid feature) The endpoint is gated by an addon the tenant has not enabled. Body carries a purchase_url. →](https://originchaindb.com/docs/examples/errors/402-addon-required)[10 403 - token valid but for a different instance The bearer is well-formed and not expired, but it does not authorise this tenant or instance. →](https://originchaindb.com/docs/examples/errors/403-forbidden) ## 404 - not found The hostname, path, or resource id does not exist. [11 404 - instance / endpoint not found The hostname or path did not match any provisioned instance or route. →](https://originchaindb.com/docs/examples/errors/404-instance)[12 404 - schema not registered The schema id in the URL is not registered for this tenant. →](https://originchaindb.com/docs/examples/errors/404-schema) ## 409 - conflict The request is valid but conflicts with current state - a concurrent writer, or an already-enabled addon. [13 409 - optimistic-concurrency conflict on row write A versioned write lost the race. Default writes are last-writer-wins; OCC is opt-in via If-Match. →](https://originchaindb.com/docs/examples/errors/409-occ-conflict)[14 409 - addon already enabled The addon-enable request hit a tenant where that addon is already on. →](https://originchaindb.com/docs/examples/errors/409-addon-already-enabled) ## 413 - payload too large The body exceeded the per-call size cap of 8 MiB. [15 413 - payload over 8 MiB The request body exceeded the per-call size limit. Split into smaller batches. →](https://originchaindb.com/docs/examples/errors/413-payload-too-large) ## 422 - semantically invalid The request parsed but failed a higher-level semantic check. [16 422 - semantically invalid manifest (FK target missing) The manifest parsed as TOML but failed semantic validation. FK targets a column that doesn't exist. →](https://originchaindb.com/docs/examples/errors/422-semantic) ## 429 - rate limited Per-API-key quota exceeded. See /docs/rate-limits for the full reference. [17 429 - rate limited (per-token cap) The per-API-key request quota was exceeded. Honour the Retry-After header before retrying. →](https://originchaindb.com/docs/examples/errors/429-rate-limited) ## 5xx - server-side The server hit an unexpected condition or the engine was momentarily unreachable. [18 500 - internal server error An unexpected server error. Trace id is included in the response; share it with support. →](https://originchaindb.com/docs/examples/errors/500-backend)[19 502 - engine unreachable (failover in progress) The service could not reach your instance. Usually transient; SDKs auto-retry. →](https://originchaindb.com/docs/examples/errors/502-engine-unreachable) --- # 400 invalid vector mode - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/400-invalid-vector-mode Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 5 / 19 # 5. 400 - invalid vector mode [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means Vector topk has exactly two modes: `fast` (approximate, IVF-style) and `high_recall` (re-rank for accuracy). Any other string returns `invalid_argument`. The handler validates the enum before touching the index. what triggers it Any mode value other than the two legal ones. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"query": [0.1, 0.2, 0.3], "k": 5, "mode": "warp_speed"}' ``` the canonical response body ``` { "error": "invalid_argument", "message": "mode 'warp_speed' is not supported; valid values are 'fast' and 'high_recall'", "retry": false } ``` how to recover - Use `fast` for online queries where p99 latency matters more than perfect recall. - Use `high_recall` for batch jobs, evals, or when you need the re-rank pass. - Omitting the field defaults to `fast`. - `retry: false` - the same payload will fail the same way. common upstream causes - Typo - `high-recall` with a hyphen, `highrecall` with no separator. - Carrying over a mode name from another vendor's API (`exact`, `flat`, `approximate`). - Configuration string read from an env var that wasn't validated at startup. - Casing - the enum is lower-case; `Fast` is rejected. --- # 400 malformed JSON - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/400-malformed-json Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 1 / 19 # 1. 400 - malformed JSON body [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The HTTP request reached the engine, but the body bytes did not parse as JSON. The router rejects before any handler runs, so no rows were touched. The error code is `invalid_json` and the message points at the byte offset where the parser gave up. what triggers it A trailing comma after the last key in the JSON body. The closing `}` arrives where the parser expected another key. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1",}' ``` the canonical response body ``` { "error": "invalid_json", "message": "trailing comma at byte offset 18", "retry": false } ``` how to recover - Read the message - it tells you the byte offset where the parse failed. - Lint your payload through a JSON validator (or your language's stdlib JSON encoder) before sending. - If you're hand-constructing JSON in a shell with quoting tricks, switch to a heredoc or a temp file and `--data-binary @body.json`. - `retry: false` - the same bytes will fail the same way. Fix the body first. common upstream causes - Trailing comma after the last field (JSON5 habit leaking into strict JSON). - Unterminated string - a missing closing `"` from shell-escape collisions. - Wrong `Content-Type` header (e.g. `text/plain`); the router treats those as invalid JSON. - Empty body on an endpoint that requires one - the parser sees EOF where it expected `{`. - String-templated payloads with unescaped quotes from user input. --- # 400 type mismatch - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/400-type-mismatch Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 4 / 19 # 4. 400 - column type mismatch [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means A value in a row write did not match the column's declared type. OriginChainDB validates writes against the manifest before they hit the storage substrate - mismatched values are rejected at the boundary, not silently coerced. The same check fires for literals on the right-hand side of a `WHERE` filter. what triggers it Putting a string into a column declared as `float64`. #### cURL ``` curl -X PUT "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/o_42" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"id": "o_42", "amount": "not-a-number", "status": "paid"}' ``` the canonical response body ``` { "error": "type_mismatch", "message": "column 'amount' is float64; value 'not-a-number' is not a number", "retry": false } ``` how to recover - Coerce client-side before sending. `parseFloat(value)`, `float(value)`, `strconv.ParseFloat`. - Confirm the column type against the manifest - `GET /v1/tenants/:t/schemas/shop.orders`. - If the column should accept text, alter the schema rather than working around it at the call site. - `retry: false` - same bytes, same answer. common upstream causes - Form fields arriving as strings from HTML inputs without client-side parsing. - `null` sent for a non-null column. - Integer overflow - sending `9999999999` into an `int32`. - Timestamp passed as Unix seconds when the column is `timestamp_ms`. - Vector dimension mismatch on an embedding column. - Boolean sent as `"true"` (string) instead of `true` (JSON literal). --- # 400 unknown column - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/400-unknown-column Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 3 / 19 # 3. 400 - unknown column [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The table resolved fine, but a column referenced in the `SELECT`, `WHERE`, `JOIN ON`, or `GROUP BY` clause is not declared on that table's manifest. The translator stops before planning. what triggers it Selecting a column that isn't in the schema. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT bogus_col FROM shop.customers LIMIT 1"}' ``` the canonical response body ``` { "error": "unknown_column", "message": "column 'bogus_col' is not declared on table 'shop.customers'", "retry": false } ``` how to recover - Fetch the manifest: `GET /v1/tenants/:t/schemas/shop.customers` - the declared columns are listed at the top. - Fix the typo, or add the column with an online-schema-change if it should exist. - Column names are case-sensitive. `CustomerID` and `customerid` are different references. - `retry: false` - same query will keep failing. common upstream causes - Typo in the column name. - SQL written against a newer schema version than the one currently registered. - Aliasing in a JOIN where the alias scope was lost. - `SELECT *` works, then a specific column was added - check spelling against the manifest. - Confusing a JSON sub-field with a top-level column (sub-fields require `json_extract(...)`). --- # 400 unknown table - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/400-unknown-table Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 2 / 19 # 2. 400 - unknown table [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The SQL parsed fine, but the table named in the `FROM` (or `JOIN`, or `UPDATE`) clause is not declared in any registered schema for this tenant. This is the SQL-layer mirror of a `schema_not_found` 404 - same root cause, different surface. what triggers it A `SELECT` against a table name that doesn't exist in the catalog. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT * FROM does_not_exist"}' ``` the canonical response body ``` { "error": "unknown_table", "message": "table 'does_not_exist' is not declared in any schema for tenant", "retry": false } ``` how to recover - List schemas for the tenant: `GET /v1/tenants/:t/schemas`. Find the one you meant. - Table names are case-sensitive and fully qualified - `shop.customers`, not `Customers`. - If the schema is missing entirely, register the manifest first via `POST /v1/tenants/:t/schemas`. - `retry: false` - the same SQL will keep failing until the schema is in place. common upstream causes - Pointing at the wrong tenant in the URL (so the schema exists, but not for this tenant). - Forgetting the schema namespace: `customers` instead of `shop.customers`. - A schema migration that renamed the table on one environment but not the other. - Copy-pasted SQL from another tenant's namespace. --- # 401 token expired - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/401-expired Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 8 / 19 # 8. 401 - token revoked or expired [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The token parsed correctly, but is either past its expiry timestamp or was revoked from the admin console. The engine treats both cases identically - the token is no longer trusted. what triggers it A bearer token whose `exp` claim is in the past, or whose ID has been added to the revocation set. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer eyJ...expired..." \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` { "error": "unauthorized", "message": "token expired at 2026-04-30T12:00:00Z", "retry": false } ``` how to recover - Issue a fresh token in the dashboard at [/app/tokens](https://app.originchain.ai/keys) and update your secret store. - If the old token was revoked (security event), rotate any downstream consumers that share it. - The SDKs surface this as a structured `TokenExpiredError` so you can wire automatic refresh in long-running services. - `retry: false` - the same token will keep failing. Refresh first. common upstream causes - Token issued with a short TTL for an integration test that ran later than expected. - System clock drift between the client and the engine - if the client clock is far ahead, valid tokens look expired. - Team member rotated the token in the dashboard but didn't redeploy the consumer. - Token explicitly revoked after an offboarding. --- # 401 malformed bearer - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/401-malformed-bearer Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 7 / 19 # 7. 401 - malformed bearer token [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The `Authorization` header arrived, but it does not match the expected shape `Bearer `. Either the `Bearer` prefix is missing, the token bytes are not a valid token, or there's stray punctuation. what triggers it An `Authorization` header that's missing the `Bearer` prefix. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: not-a-real-bearer" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` { "error": "unauthorized", "message": "Authorization header must be 'Bearer '", "retry": false } ``` how to recover - Prefix the token with the literal word `Bearer` (capital B, single space). - Echo the env var to confirm it has the actual token bytes, not `"$OC_TOKEN"` as a literal string. - Strip newlines from a token pasted out of a wrapped terminal - some shells inject `\\r`. - `retry: false` - the header is wrong, not the engine. common upstream causes - Token sent without the `Bearer` prefix. - Lowercase `bearer` from a snippet - the parser is case-sensitive on the scheme name. - Smart-quote characters from a doc copy-paste (`"` vs `"`). - Trailing whitespace or newline pulled in from a file read. - Two tokens concatenated by accident (env var set twice). --- # 401 missing bearer - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/401-missing-bearer Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 6 / 19 # 6. 401 - missing Authorization header [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means No `Authorization` header was sent. The router rejects unauthenticated traffic before any handler runs. Every API call requires `Authorization: Bearer `. what triggers it A request that omits the `Authorization` header entirely. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` { "error": "unauthorized", "message": "missing Authorization header", "retry": false } ``` how to recover - Add `-H "Authorization: Bearer $OC_TOKEN"` to the request. - Issue a token in the dashboard ([/app/tokens](https://app.originchain.ai/keys)) if you don't have one yet. - If you're using one of the SDKs, set `OC_TOKEN` in the environment - the clients pick it up automatically. - `retry: false` - the call needs a header, not a backoff. common upstream causes - Token env var unset in CI - the script silently sent an empty value, which the proxy stripped. - A reverse proxy or sidecar dropping the `Authorization` header on forward. - Calling through a CDN cache that wasn't configured to forward auth headers. - Pasted cURL from a docs page that didn't include the auth line. - Frontend code that forgot to attach the token to `fetch` options. --- # 402 addon required - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/402-addon-required Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 9 / 19 # 9. 402 - addon required [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The endpoint is gated by an addon the tenant has not enabled yet. Unlike most error bodies, a 402 carries extra typed fields: the addon's short `code`, its human `name`, the `monthly_usd` price, and a `purchase_url` the dashboard surfaces as a one-click enable. what triggers it Calling a paid endpoint (here, natural-language SQL via `/ask`) before the `ask` addon is enabled. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"question": "how many customers do I have?"}' ``` the canonical response body ``` { "error": "addon_required", "message": "natural-language SQL requires the 'ask' addon", "retry": false, "addon": "ask", "name": "Natural language SQL", "monthly_usd": 99, "purchase_url": "https://originchaindb.com/app/addons?enable=ask" } ``` how to recover - Open the `purchase_url` from the response body - the dashboard turns it into a one-click enable, no support ticket required. - Or enable programmatically: `POST /v1/tenants/:t/addons` with `{"addon": "ask"}`. - If you're building a UI on top of OriginChainDB, surface the `name` and `monthly_usd` fields so the user sees what they're enabling. - `retry: false` - the call needs the addon, not a retry. Once enabled, the same request will succeed. common upstream causes - New tenant exploring paid endpoints during onboarding. - Different tenant in the same org has the addon - tokens are per-instance, not org-wide. - Free-configuration customer hitting a paid surface from a tutorial. --- # 403 forbidden - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/403-forbidden Sitemap last modified: 2026-09-14T18:28:08.000Z examples · errors · 10 / 19 # 10. 403 - token valid but for a different instance [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The bearer is well-formed and not expired, but it does not authorise the tenant in the URL. OriginChainDB tokens are scoped to a single instance - on a dedicated configuration, sending one tenant's token to another tenant's endpoint returns 403 and never reveals whether the other tenant exists. what triggers it A token for tenant A used against tenant B's URL. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/other-tenant/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` { "error": "forbidden", "message": "token does not authorise access to tenant 'other-tenant'", "retry": false } ``` how to recover - Confirm the tenant id in the URL matches the tenant the token was issued for. - If you genuinely need to call the other tenant, issue a token from that tenant's dashboard. - Multi-tenant apps should key the token off the tenant before constructing the URL - never the other way around. - `retry: false` - same token, same URL, same answer. common upstream causes - Tenant id hardcoded in a snippet that was copied between environments. - Staging token used against a production endpoint. - Tenant id and instance hostname out of sync after a rename. - Multi-tenant app router building the URL from one source and the token from another. - Token scope was narrowed (e.g. read-only) and the call needs a permission the scope doesn't grant. --- # 404 instance not found - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/404-instance Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 11 / 19 # 11. 404 - instance / endpoint not found [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The hostname resolved but didn't match any provisioned OriginChainDB instance, or the path is not a registered route. The router replies with a generic `not_found` rather than disclosing what does exist. what triggers it A request to a hostname that doesn't belong to any tenant. #### cURL ``` curl -X POST "https://nope.db.originchain.ai/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` { "error": "not_found", "message": "no instance is provisioned at this endpoint", "retry": false } ``` how to recover - Confirm the instance hostname in the dashboard - it's on the instance details page. - Check the path is one of the documented routes (`/sql`, `/rows`, `/vector`, `/graph`, `/ask`, `/schemas`). - If the instance was destroyed, provision a new one or point at a sibling instance. - `retry: false` - the URL is wrong, not the engine. common upstream causes - Typo in the instance subdomain. - Typo in the tenant subdomain (the part before `.db.originchain.ai`). - Instance was destroyed; the env var still points at the old endpoint. - Pre-launch hostname queued in DNS but provisioning hasn't completed yet. - Calling `/api/v1` instead of `/v1`, or a roadmap endpoint that hasn't shipped. --- # 404 schema not found - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/404-schema Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 12 / 19 # 12. 404 - schema not registered [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The URL targets a schema id that is not registered for this tenant. Same root cause as the SQL `unknown_table` 400, but surfaced at the REST resource layer rather than through the SQL translator. what triggers it A direct GET against a schema id that doesn't exist. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/schemas/does.not.exist" \ -H "Authorization: Bearer $OC_TOKEN" ``` the canonical response body ``` { "error": "schema_not_found", "message": "no schema 'does.not.exist' for tenant", "retry": false } ``` how to recover - List schemas registered for this tenant: `GET /v1/tenants/:t/schemas`. - Register the manifest with `POST /v1/tenants/:t/schemas` if it should exist. - Confirm the schema id is fully qualified - the catalog stores `shop.customers`, not `customers`. - `retry: false` - register the schema first, then retry. common upstream causes - Schema registered on staging but not yet on prod. - Typo in the namespace - `shop.` vs `shops.`. - Schema was deleted (via a destructive-data confirm) and a downstream service still references it. - Multi-region deploy where the new region missed a migration step. --- # 409 addon already enabled - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/409-addon-already-enabled Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 14 / 19 # 14. 409 - addon already enabled [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The request tried to enable an addon that is already enabled for this tenant. The endpoint is intentionally not idempotent - it returns 409 so the caller knows the state didn't change. what triggers it Re-enabling an addon that's already on. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/addons" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"addon": "ask"}' ``` the canonical response body ``` { "error": "addon_already_enabled", "message": "addon 'ask' is already enabled for this tenant", "retry": false } ``` how to recover - Check the current state first: `GET /v1/tenants/:t/addons` returns the list of enabled addons. - If your script is meant to be idempotent, treat this 409 as a no-op success. - If you wanted a different addon, fix the body and retry with the right code. - `retry: false` - the state is already what you wanted. common upstream causes - Provisioning script run twice in CI. - Dashboard click that re-fired through a stale tab. - Migration that enabled the addon manually, then a Terraform run that tried to enable it again. - Two team members enabling the same addon from different sessions. --- # 409 OCC conflict - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/409-occ-conflict Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 13 / 19 # 13. 409 - optimistic-concurrency conflict [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means By default, OriginChainDB row writes are last-writer-wins - two concurrent `PUT`s on the same key both succeed, and the later one overwrites the earlier one. No 409. OCC (optimistic concurrency control) is opt-in. You enable it by passing the row's current version in an `If-Match: ` header. If the row's actual version differs at commit time, the write is rejected with this 409 and the body tells you the current version so you can re-fetch, re-apply, and retry. what triggers it A conditional PUT where the `If-Match` version no longer matches. #### cURL ``` curl -X PUT "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.customers/c_42" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "If-Match: 17" \ -H "Content-Type: application/json" \ -d '{"id": "c_42", "email": "alice@example.com", "tier": "platinum"}' ``` the canonical response body ``` { "error": "write_conflict", "message": "row 'c_42' is at version 18, expected 17", "retry": true, "current_version": 18 } ``` how to recover - Re-read the row to get the current state and version. - Re-apply your business logic on top of the new state (merge, re-validate, etc). - Retry the PUT with `If-Match: `. - The SDKs offer a `update_with_retry()` helper that wraps the read-merge-retry loop. - `retry: true` - but only after re-fetching. Blind retry of the same bytes will fail the same way. common upstream causes - Two workers picking up the same job from a queue and racing on the same row. - Background reconciler and user-facing edit landing in the same millisecond. - Re-using a stale version number from an in-memory cache. - Mobile client with an offline edit syncing after the server already advanced. - Wrong row's version copied into `If-Match` from another response. --- # 413 payload too large - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/413-payload-too-large Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 15 / 19 # 15. 413 - payload over 8 MiB [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The request body exceeded the per-call cap of 8 MiB. The router rejects oversize bodies at the edge so they never touch the engine. The cap applies to the raw bytes after any compression at the HTTP layer is reversed. what triggers it A bulk insert from a file larger than 8 MiB. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.products/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ --data-binary @huge-batch.json ``` the canonical response body ``` { "error": "payload_too_large", "message": "request body 12.4 MiB exceeds limit of 8 MiB", "retry": false, "limit_bytes": 8388608 } ``` how to recover - Split the batch. The SDKs' `insert_many()` helpers chunk automatically; if you're calling REST directly, chunk into <= 5 MiB groups. - For ongoing high-throughput ingestion, use the streaming ingest endpoints rather than bulk POSTs. - Trim large blob fields from the row payload - if you're ingesting big binaries, put them in object storage and store a reference key. - `retry: false` - same body, same limit. Smaller batches will succeed. common upstream causes - A single batch insert containing every record from a CSV dump. - Rows that include large base64-encoded blobs (images, PDFs). - Wide vector embeddings on many rows (768-dim float32 = 3 KiB per row before JSON overhead). - A debug payload that accidentally inlined a stack trace or a large log fragment. - Forgetting to paginate an upstream source. --- # 422 invalid manifest - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/422-semantic Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 16 / 19 # 16. 422 - semantically invalid manifest [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The request parsed correctly (TOML is syntactically valid) but failed a higher-level semantic check. 422 is the bucket for "syntax fine, meaning wrong" - foreign key targets that don't exist, primary keys with no columns, vector dims that don't match the declared embedding model, etc. what triggers it A manifest where a foreign key points at a column that doesn't exist on the target table. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/schemas" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/toml" \ --data-binary @- <<'TOML' [table] name = "shop.orders" [[column]] name = "id" type = "string" primary_key = true [[column]] name = "customer_id" type = "string" [[foreign_key]] column = "customer_id" references = "shop.customers.missing_pk" TOML ``` the canonical response body ``` { "error": "invalid_manifest", "message": "foreign key target 'shop.customers.missing_pk' does not exist", "retry": false } ``` how to recover - Read the message - it names the exact target the validator could not find. - Fetch the referenced table's manifest and confirm the column exists with the spelling you used. - If the target table needs to be created or extended first, do that before registering this manifest. - `retry: false` - same bytes, same answer. Fix the manifest first. common upstream causes - FK target table not registered yet (apply manifests in the right order). - Typo in the FK target column name. - Primary key column missing the `primary_key = true` flag on the target. - Vector column's declared dim doesn't match the embedding model. - Two indexes with the same name on one table. - CHECK constraint with the field `expr` (legacy) instead of the current `expression`. --- # 429 rate limited - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/429-rate-limited Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 17 / 19 # 17. 429 - rate limited (per-token cap) [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The API key has exceeded its per-second request quota. Limits are per token, not per instance - if you have several services sharing one token, they share the budget. The response includes both the standard `Retry-After` header and a `retry_after_seconds` body field. See [/docs/rate-limits](https://originchaindb.com/docs/rate-limits) for the full reference (per-configuration caps, burst behaviour, headers). what triggers it Any request after the per-token cap is hit. `-i` shows the headers. #### cURL ``` curl -i -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` HTTP/1.1 429 Too Many Requests Retry-After: 7 Content-Type: application/json { "error": "rate_limited", "message": "API key has exceeded its per-second request quota", "retry": true, "retry_after_seconds": 7 } ``` how to recover - Sleep for `Retry-After` seconds before retrying. Don't tight-loop. - The SDKs honour `Retry-After` automatically with exponential backoff and jitter. - If you're consistently hitting the cap, raise your configuration (a larger configuration has a higher per-token quota) or split traffic across multiple tokens. - Audit for accidental hot loops - a worker that retries with no delay stays pinned at the cap. - `retry: true` - after the indicated wait. common upstream causes - Backfill script with no per-request delay. - One token shared across many services that each assume they have the full budget. - A retry loop that doesn't honour `Retry-After` and amplifies the spike. - Multiple replicas of a worker that all woke up after a deploy and started together. - A frontend that fires a fetch per keystroke instead of debouncing. --- # 500 internal server error - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/500-backend Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 18 / 19 # 18. 500 - internal server error [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The server hit an unexpected condition. We log every 500 internally; the response body and the `x-oc-trace-id` response header both carry the trace id so support can find the exact request. The body also carries a `status_page` URL - check there first to see if it's a known incident. what triggers it Any normally-valid request that surfaces a bug or unhandled internal condition. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` { "error": "internal_error", "message": "unexpected backend error; please retry. If this persists, contact support with the trace id below.", "retry": true, "trace_id": "0b1c2d3e-4f50-6172-8390-a1b2c3d4e5f6", "status_page": "https://status.originchain.ai" } ``` how to recover - Check [status.originchain.ai](https://status.originchain.ai/). If there's an open incident, the ETA is there. - Retry with exponential backoff - 500s are often transient. - If it persists, open a support ticket and include the `trace_id`. With that id we can pull the exact request out of the logs. - Don't assume the write didn't happen - 500 means we don't know. Use a read or OCC retry to confirm state before re-issuing a non-idempotent write. - `retry: true` - with backoff. common upstream causes - A new bug surfaced by a recent deploy. - Engine OOM on an unusually large query. - Underlying storage substrate I/O error. - A panic in an edge code path (logged, paged, fixed forward). - An internal dependency (the platform or billing service) was momentarily down. --- # 502 engine unreachable - OriginChainDB error example Canonical source: https://originchaindb.com/docs/examples/errors/502-engine-unreachable Sitemap last modified: 2026-09-13T20:18:50.000Z examples · errors · 19 / 19 # 19. 502 - engine unreachable [← Errors examples](https://originchaindb.com/docs/examples/errors) what this error means The service received the request but could not reach your instance. It is almost always transient - typically your instance is briefly unreachable while it restarts: a planned config change, a rolling engine deploy, or a writer changeover in an active-passive pair. The official SDKs auto-retry on 502 with exponential backoff. If you're calling REST directly, you should too. what triggers it Any request that lands in that window - a restart, a deploy, or a writer changeover. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql": "SELECT id FROM shop.products LIMIT 1"}' ``` the canonical response body ``` { "error": "engine_unreachable", "message": "control plane could not reach the engine for this tenant", "retry": true, "trace_id": "5f4e3d2c-1b0a-9988-7766-554433221100", "status_page": "https://status.originchain.ai" } ``` how to recover - Retry with exponential backoff. These windows are normally short; anything longer will show on the status page. - If you're on an SDK, the retry happens automatically - you'll usually only see the 502 in logs as a warning, not as a thrown error. - Check [status.originchain.ai](https://status.originchain.ai/) if retries keep failing for more than a minute. - For non-idempotent writes, prefer OCC or a deduping pattern so a retried request doesn't double-apply. - `retry: true` - this is the canonical "retry me" error. common upstream causes - A writer changeover in an active-passive pair - a standby takes over as writer. Promotion is operator-driven by default; the engine also has an opt-in automatic promotion path, off by default, which depends on an external restart hook and refuses to promote a standby that is not fully caught up. - Planned restart for a config change. - A transient network blip inside the service. - Engine briefly unresponsive during a heavy compaction. - A rolling deploy of the engine binary. --- # Full-text search examples - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/fts Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts # Full-text search examples [← All examples](https://originchaindb.com/docs/examples) FTS on OriginChainDB is a runtime endpoint: you POST documents to a field and query them back, and no schema declaration is required. Nine examples follow, each on its own page with side-by-side cURL / Python / TypeScript / Go, the response shape, and notes on common mistakes. newer search capabilities The native endpoint on this page covers BM25, boolean and phrase search. For fuzzy matching (typo tolerance), per-field analyzers with language stemming, and multi-field / boosted queries, use the [Elasticsearch-compatible API](https://originchaindb.com/docs/examples/elasticsearch) — a drop-in for the `@elastic` clients that runs over these same substrate-native indexes. [Connect a client →](https://originchaindb.com/docs/connect-elasticsearch) The endpoint takes the form `/v1/tenants/:t/fts//`. Each field is its own independent inverted index. The `doc_id` you supply at index time is what comes back at search time - typically the row's primary key. [1 Index a document (write postings) POST a doc_id + text to the FTS endpoint. Re-indexing the same doc_id atomically replaces the previous text. →](https://originchaindb.com/docs/examples/fts/index-doc)[2 Boolean search - single token Ask whether any doc contains a token. Returns matching doc_ids only, no ranking. →](https://originchaindb.com/docs/examples/fts/boolean-single)[3 Boolean search - multi-token AND All tokens must match. Returns doc_ids. There is no OR in boolean mode - union client-side if you need it. →](https://originchaindb.com/docs/examples/fts/boolean-and)[4 Phrase search Tokens must appear in order, contiguously. Use for branded phrases, model numbers, log templates. →](https://originchaindb.com/docs/examples/fts/phrase)[5 BM25 ranked retrieval Lucene-default scoring (k1=1.2, b=0.75). Returns { doc_id, score } sorted by relevance. Use when ordering matters. →](https://originchaindb.com/docs/examples/fts/bm25)[6 Multi-field weighted merge Index two fields separately, query each, merge results client-side with your own weights. →](https://originchaindb.com/docs/examples/fts/multi-field)[7 Aggregations over a result set Bucket and measure the documents a query matched - terms, range, stats and cardinality folded over the whole match set, never the top-k. →](https://originchaindb.com/docs/examples/fts/aggregations)[8 Autocomplete and did-you-mean Prefix completion, per-term spelling correction and whole-phrase rewriting from one endpoint, selected by the kind field. →](https://originchaindb.com/docs/examples/fts/suggest)[9 Synonyms and stopwords Make tv and television interchangeable, or replace the built-in stopword list for one field. Both replace the stored record in full. →](https://originchaindb.com/docs/examples/fts/synonyms-stopwords) --- # Aggregations over a result set - OriginChainDB FTS Canonical source: https://originchaindb.com/docs/examples/fts/aggregations Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 7 / 9 # 7. Aggregations over a result set [← FTS examples](https://originchaindb.com/docs/examples/fts) ## what this does Folds an aggregation tree over the documents a query matched and returns bucket counts and metrics instead of hits - "what is in this result set" rather than "what ranks highest in it". The values come from per-document records you write separately, so you can group by an attribute that was never part of the indexed text. ## when to use it - Faceted navigation - the colour, brand and price-band counts shown beside a result list. - Dashboards over a filtered slice: revenue, distinct values, min / max of a numeric attribute. - Any question whose answer is a number about the whole match set rather than a page of documents. ## write the values first An aggregation reads per-document records, not the indexed text. Postings written by [Example 1](https://originchaindb.com/docs/examples/fts/index-doc) carry no attribute values, so give each document a `facets` map first. Re-posting the same `doc_id` replaces its record. ``` # Attach the values the aggregation will fold over. "text" is the stored # text highlighting reads; "facets" maps a facet field to its values. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description/doc" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "p001", "text": "Over-ear headphones with active noise cancellation", "facets": { "color": ["red"], "price": ["249"] } }' ``` ## the request Bucket every matching document by `color` and sum `price` inside each bucket. Omit `search` entirely and the tree runs over every document indexed under the field. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/_aggs" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "field": "description", "search": { "query": { "bool": { "must": [{ "match": { "field": "description", "query": "headphones" } }], "must_not": [{ "term": { "field": "description", "value": "refurbished" } }] } } }, "aggs": { "by_color": { "terms": { "field": "color" }, "aggs": { "revenue": { "sum": { "field": "price" } } } } } }' ``` #### Python ``` import requests # _aggs hangs off the SCHEMA, not the field. "field" in the body names the # FTS field whose per-document records supply the values to aggregate. resp = requests.post( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/fts/shop.products/_aggs", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "field": "description", "search": {"query": {"bool": { "must": [{"match": {"field": "description", "query": "headphones"}}], "must_not": [{"term": {"field": "description", "value": "refurbished"}}], }}}, "aggs": { "by_color": { "terms": {"field": "color"}, "aggs": {"revenue": {"sum": {"field": "price"}}}, } }, }, ) result = resp.json() for b in result["aggs"]["by_color"]["buckets"]: print(b["key"], b["doc_count"], b["aggs"]["revenue"]["value"]) ``` #### TypeScript ``` // _aggs hangs off the schema; "field" names the FTS field whose // per-document records supply the values to aggregate. const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/fts/shop.products/_aggs`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ field: "description", search: { query: { bool: { must: [{ match: { field: "description", query: "headphones" } }], must_not: [{ term: { field: "description", value: "refurbished" } }], } } }, aggs: { by_color: { terms: { field: "color" }, aggs: { revenue: { sum: { field: "price" } } }, }, }, }), }, ); const result = await res.json(); for (const b of result.aggs.by_color.buckets) { console.log(b.key, b.doc_count, b.aggs.revenue.value); } ``` #### Go ``` // _aggs hangs off the schema; "field" names the FTS field whose // per-document records supply the values to aggregate. body, _ := json.Marshal(map[string]any{ "field": "description", "search": map[string]any{"query": map[string]any{"bool": map[string]any{ "must": []any{map[string]any{ "match": map[string]any{"field": "description", "query": "headphones"}, }}, "must_not": []any{map[string]any{ "term": map[string]any{"field": "description", "value": "refurbished"}, }}, }}}, "aggs": map[string]any{ "by_color": map[string]any{ "terms": map[string]any{"field": "color"}, "aggs": map[string]any{"revenue": map[string]any{"sum": map[string]any{"field": "price"}}}, }, }, }) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, "https://"+ocHost+"/v1/tenants/"+ocTenant+"/fts/shop.products/_aggs", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+ocToken) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` { "doc_count": 2, "aggs": { "by_color": { "kind": "buckets", "buckets": [ { "key": "red", "doc_count": 2, "aggs": { "revenue": { "kind": "value", "value": 349.0 } } } ], "sum_other_doc_count": 0 } } } ``` `doc_count` is the size of the set the tree was evaluated over. Every result is tagged with `kind` - `buckets`, `value`, `stats`, `cardinality` and so on - so a client can dispatch on the shape without knowing which aggregation produced it. ## how it works - Two different "field" keys. The top-level `field` names the FTS field whose per-document records hold the values. The `field` inside each aggregation names the facet being bucketed. They are rarely the same string. - Buckets cover the whole match set. Ranked search stops at a top-k; an aggregation cannot, or it would answer a different question with a number that looks right. The executor is asked for every match, so a `terms` count counts all of them. - Too many matches is a refusal, not a sample. A match set above the hit ceiling (`OC_FTS_MAX_HITS`, default 10,000) returns 400 naming the limit rather than quietly aggregating a prefix. - Trees nest. Each bucket carries its own `aggs` map, up to `MAX_AGG_DEPTH` = 8 levels. `terms` buckets default to count-descending. - Facet values are stored as strings and coerced to numbers by the numeric aggregations, which is why `"249"` sums correctly. ## common mistakes - Aggregating a field that has no per-document records. Indexing text writes postings only. Without a `/doc` write carrying `facets`, there is nothing to fold and the bucket lists come back empty. - Sending an empty `aggs` object. Refused with 400 - an empty tree would return an object that reads like a result. - A pipeline aggregation at the top level. `cumulative_sum` and its siblings act on a parent bucket list; at the root there is none, so the request is refused rather than answered with zeroes. - Expecting `top_k` to bound the cost. It is ignored here by design. If the query matches too much, narrow the query - the endpoint will not hand back a prefix. - Putting `doc_values_field` in the `search` block. Refused: the aggregated `field` is already the value source, and the two could only disagree. --- # BM25 ranked retrieval - OriginChainDB FTS example Canonical source: https://originchaindb.com/docs/examples/fts/bm25 Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 5 / 9 # 5. BM25 ranked retrieval [← FTS examples](https://originchaindb.com/docs/examples/fts) what this does Returns the top `k` documents ranked by BM25 relevance to the query. Plain BM25 returns a bare JSON array - each entry is `{ doc_id, score }`, sorted by descending score. Higher score = more relevant. (Boolean and phrase modes return arrays of plain doc_id strings; BM25 adds the score.) when to use it - Search bars - users expect the best match at the top. - Recommendations: rank a candidate pool by textual similarity. - Any time you want "best k" rather than "all that match". the request Assumes you have indexed at least a few documents - see [Example 1](https://originchaindb.com/docs/examples/fts/index-doc). #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=wireless headphones" \ --data-urlencode "mode=bm25" \ --data-urlencode "k=10" ``` #### Python ``` # Plain BM25 returns a bare JSON array of { doc_id, score } objects hits = db.fts.search( "shop.products", "description", q="wireless headphones", mode="bm25", k=10, ) for hit in hits: print(hit["doc_id"], hit["score"]) ``` #### TypeScript ``` // Plain BM25 returns a bare array of { doc_id, score } const hits = await db.ftsSearch("shop.products", "description", { q: "wireless headphones", mode: "bm25", k: 10, }); for (const hit of hits) { console.log(hit.doc_id, hit.score); } ``` #### Go ``` // Plain BM25 returns []{ DocID, Score } hits, _ := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "wireless headphones", Mode: "bm25", K: 10, }) for _, hit := range hits { fmt.Println(hit.DocID, hit.Score) } ``` what you get back ``` [ { "doc_id": "p001", "score": 9.42 }, { "doc_id": "p027", "score": 7.18 }, { "doc_id": "p014", "score": 4.55 } ] ``` A bare array of `{ doc_id, score }` objects - no `{ "mode": ..., "hits": [...] }` wrapper. the enriched shape (highlight / facets only) The plain bare-array shape only changes when you add `highlight=true` or `facets=`. Then - and only then - BM25 returns an object with a `hits` array (each hit carrying optional `highlights`) and a `facets` block. Highlights require the doc text to have been stored via `POST /fts/:t/:f/doc` first. ``` { "hits": [ { "doc_id": "p001", "score": 9.42, "highlights": { "description": ["Over-ear wireless headphones with…"] } }, { "doc_id": "p027", "score": 7.18, "highlights": {} } ], "facets": { "category": [{ "value": "electronics", "count": 2 }] } } ``` `k`, `fuzzy`, `highlight`, `facets`, and `explain` are BM25-only - they are silently ignored in boolean and phrase modes. how it works - The query is tokenised, then for each token the engine fetches the posting list along with term frequencies and document lengths. - Each candidate document gets a BM25 score using the Lucene defaults: `k1 = 1.2`, `b = 0.75`. Score grows with how often the rare terms appear and shrinks if the document is much longer than average. - A top-`k` heap keeps the highest scorers; everything else is discarded. Optional query params: `fuzzy=1` for single-character typo tolerance, `highlight=true` to return matched-snippet text, `facets=col,col` for grouped counts alongside hits. common mistakes - Comparing scores across queries. A BM25 score of 9.4 means nothing on its own and can't be compared to the 9.4 from a different query. Use scores to rank within one result set only. - Forgetting `k`. If you omit it, you get a default cap. Set `k` to what you actually need; ranking the entire corpus is wasted work. - Reaching for BM25 when boolean would do. If you only need "does it match", boolean is cheaper and the answer is the same set. --- # Boolean search - multi-token AND - OriginChainDB FTS example Canonical source: https://originchaindb.com/docs/examples/fts/boolean-and Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 3 / 9 # 3. Boolean search - multi-token AND [← FTS examples](https://originchaindb.com/docs/examples/fts) what this does Returns documents whose `description` contains both `wireless` and `headphones`. Tokens are intersected. The tokens can appear in any order and anywhere in the document - for "in order, next to each other," use [phrase mode](https://originchaindb.com/docs/examples/fts/phrase). when to use it - Filtering down a result set by multiple required attributes ("organic" AND "coffee"). - Tag intersections - any doc that mentions every tag. - You don't care about ranking and you want the smallest possible response. the request Tokens are space-separated in `q`. Some clients show this URL-encoded as `wireless+headphones` - same thing. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=wireless headphones" \ --data-urlencode "mode=boolean" ``` #### Python ``` # Boolean search returns a bare JSON array of doc_id strings doc_ids = db.fts.search( "shop.products", "description", q="wireless headphones", mode="boolean", ) for doc_id in doc_ids: print(doc_id) ``` #### TypeScript ``` // Boolean search returns a bare string[] of doc_ids const docIds = await db.ftsSearch("shop.products", "description", { q: "wireless headphones", mode: "boolean", }); for (const docId of docIds) { console.log(docId); } ``` #### Go ``` // Boolean search returns []string of doc_ids docIDs, _ := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "wireless headphones", Mode: "boolean", }) for _, docID := range docIDs { fmt.Println(docID) } ``` what you get back ``` ["p001", "p027"] ``` Bare array of `doc_id` strings (the intersection), sorted lexicographically. No wrapper object, no scores. how it works - The query is tokenised into a list - here `["wireless", "headphones"]`. - The engine fetches the posting list for each token and intersects them. Smallest list first keeps the work bounded by the rarest term. - The intersection is the answer. No scoring is computed. common mistakes - Assuming OR behaviour. Multi-token boolean is strictly AND - there is no OR mode. For OR, run two separate queries and union the returned arrays client-side. - Expecting proximity. AND doesn't care where the tokens sit - "wireless mouse and wired headphones" matches "wireless headphones". Use phrase mode if you need them adjacent. - Adding a rare token to "help match more". AND only narrows - every extra token can only remove docs, never add. --- # Boolean search - single token - OriginChainDB FTS example Canonical source: https://originchaindb.com/docs/examples/fts/boolean-single Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 2 / 9 # 2. Boolean search - single token [← FTS examples](https://originchaindb.com/docs/examples/fts) what this does Returns the `doc_id` of every document whose `description` contains the token `wireless`, as a bare JSON array of strings. No scoring is computed and no ordering is implied - this is a set, not a ranked list. Boolean is the default mode; you get it even if you omit `mode`. when to use it - "Does any product mention X" - membership checks, not relevance. - Pre-filter step before a more expensive query (rank only the matched set). - Tag-style search where every match is equally important. If you need ordering by relevance, use [BM25 mode](https://originchaindb.com/docs/examples/fts/bm25) instead. the request Assumes you've indexed at least a few documents - see [Example 1](https://originchaindb.com/docs/examples/fts/index-doc). #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=wireless" \ --data-urlencode "mode=boolean" ``` #### Python ``` # Boolean search returns a bare JSON array of doc_id strings doc_ids = db.fts.search( "shop.products", "description", q="wireless", mode="boolean", ) for doc_id in doc_ids: print(doc_id) ``` #### TypeScript ``` // Boolean search returns a bare string[] of doc_ids const docIds = await db.ftsSearch("shop.products", "description", { q: "wireless", mode: "boolean", }); for (const docId of docIds) { console.log(docId); } ``` #### Go ``` // Boolean search returns []string of doc_ids docIDs, _ := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "wireless", Mode: "boolean", }) for _, docID := range docIDs { fmt.Println(docID) } ``` what you get back ``` ["p001", "p014", "p027"] ``` The body is the array itself - there is no `{ "mode": ..., "doc_ids": [...] }` wrapper. Parse it and iterate directly. how it works - The query is tokenised with the same analyser used at index time - lowercased and split on whitespace/punctuation. - The single token is looked up in the inverted index. The posting list for that token is the answer. - Because there's no scoring step, boolean mode is the cheapest query shape - constant work per matching doc, no per-doc maths. common mistakes - Expecting a ranked order. The returned array is sorted lexicographically by `doc_id`, not by relevance. Don't read any ranking into the order. - Looking for a `.doc_ids` field. The response is the bare array - `["p001", "p014"]` - not an object. Accessing `.doc_ids` on it returns undefined. - Expecting case sensitivity. The query token is lowercased to match indexing - `q=Wireless` and `q=wireless` behave identically. - Passing multiple tokens and expecting OR. Multi-token boolean is AND - see the next example. --- # Index a document - OriginChainDB FTS example Canonical source: https://originchaindb.com/docs/examples/fts/index-doc Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 1 / 9 # 1. Index a document (write postings) [← FTS examples](https://originchaindb.com/docs/examples/fts) what this does Tokenises `text` and writes the resulting postings into the inverted index for the `description` field on `shop.products`. The `doc_id` you supply is the handle the search side returns - so it should be whatever you use to look the row up later (almost always the row's primary key). when to use it - You have a text column (description, body, comment) and you want to be able to search inside it. - You're rebuilding an index after a backfill or a schema change. - You want to update what's indexed for one row - re-POST the same `doc_id` and the previous text is replaced atomically. the request FTS is a runtime endpoint - there is no `[[extractions.fts]]` block on the schema and no setup step. The first POST creates the index. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "p001", "text": "Wireless over-ear headphones with active noise cancellation and 30-hour battery life." }' ``` #### Python ``` db.fts.index( "shop.products", "description", doc_id="p001", text="Wireless over-ear headphones with active noise cancellation and 30-hour battery life.", ) ``` #### TypeScript ``` await db.ftsIndex("shop.products", "description", { doc_id: "p001", text: "Wireless over-ear headphones with active noise cancellation and 30-hour battery life.", }); ``` #### Go ``` _, err := db.FTSIndex(ctx, "shop.products", "description", originchain.FTSIndexRequest{ DocID: "p001", Text: "Wireless over-ear headphones with active noise cancellation and 30-hour battery life.", }) ``` what you get back ``` HTTP/1.1 201 Created (empty body) ``` A successful index returns `201 Created` with no response body. There is no `tokens_indexed` count or confirmation JSON - check the status code, not the body. how it works - The text is tokenised, lowercased, and the per-token postings are written to the inverted index keyed by `(schema, field, token)`. - If a doc with this `doc_id` already exists in this field's index, the previous postings are removed and the new ones written in a single atomic batch. There is no window where both versions are visible. - Each field is an independent index. Indexing `description` does nothing to `title`. - `:table` and `:field` are opaque path segments - they are not validated against any schema. They must match exactly (byte-for-byte) between the index POST and the search GET. Index under `shop.products` and query `shop_products` and you get a silent empty result, not an error. common mistakes - Forgetting to re-index on writes. Updating the source row's text in `/rows` does not update the FTS index. You have to POST to the FTS endpoint again, or wire a hook that does it for you. - Using something other than the primary key for `doc_id`. Search returns `doc_id`s - if it's not the row's PK you'll need an extra lookup on every hit. - Indexing multi-MB blobs as a single doc. Very large documents bloat the postings and hurt scoring. Split into chunks (e.g. one doc_id per paragraph or per page) and let the search side merge. --- # Multi-field weighted merge - OriginChainDB FTS example Canonical source: https://originchaindb.com/docs/examples/fts/multi-field Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 6 / 9 # 6. Multi-field weighted merge [← FTS examples](https://originchaindb.com/docs/examples/fts) what this does Searches the same query across two fields - typically a short, high-signal field like `title` and a longer body field like `description` - and produces a single ranked list. Each field is an independent inverted index, so you make two queries and combine the scores in your app with whatever weights you want (e.g. `2.0 × title + 1.0 × description`). when to use it - Catalogue / product search where a hit in the title is worth more than a hit in the body. - Help-centre search across `heading` + `body`. - Anywhere you'd reach for "field boosts" in a traditional search engine. the request Two POSTs to set up (one per field), two GETs to query, one client-side merge step. The `doc_id` on both fields is the same row PK so the merge can join on it. #### cURL ``` # Step 1: index the same row into two separate fields curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/title" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "p001", "text": "Wireless Headphones - Pro" }' curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "p001", "text": "Over-ear wireless headphones with noise cancellation." }' # Step 2: query each field, then merge client-side curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/title" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=wireless headphones" --data-urlencode "mode=bm25" --data-urlencode "k=20" curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=wireless headphones" --data-urlencode "mode=bm25" --data-urlencode "k=20" ``` #### Python ``` # Index the row's title and description as two separate FTS docs db.fts.index("shop.products", "title", doc_id="p001", text="Wireless Headphones - Pro") db.fts.index("shop.products", "description", doc_id="p001", text="Over-ear wireless headphones with noise cancellation.") # Query each field independently title_hits = db.fts.search("shop.products", "title", q="wireless headphones", mode="bm25", k=20) desc_hits = db.fts.search("shop.products", "description", q="wireless headphones", mode="bm25", k=20) # Each plain-BM25 call returns a bare array of { doc_id, score } — iterate directly def merge(title_hits, desc_hits, w_title=2.0, w_desc=1.0): scores: dict[str, float] = {} for h in title_hits: scores[h["doc_id"]] = scores.get(h["doc_id"], 0.0) + w_title * h["score"] for h in desc_hits: scores[h["doc_id"]] = scores.get(h["doc_id"], 0.0) + w_desc * h["score"] return sorted(scores.items(), key=lambda kv: kv[1], reverse=True) for doc_id, score in merge(title_hits, desc_hits)[:10]: print(doc_id, score) ``` #### TypeScript ``` // Index the row's title and description as two separate FTS docs await db.ftsIndex("shop.products", "title", { doc_id: "p001", text: "Wireless Headphones - Pro" }); await db.ftsIndex("shop.products", "description", { doc_id: "p001", text: "Over-ear wireless headphones with noise cancellation." }); // Query each field independently const titleHits = await db.ftsSearch("shop.products", "title", { q: "wireless headphones", mode: "bm25", k: 20 }); const descHits = await db.ftsSearch("shop.products", "description", { q: "wireless headphones", mode: "bm25", k: 20 }); // Merge client-side with your own weights function merge( title: typeof titleHits, desc: typeof descHits, wTitle = 2.0, wDesc = 1.0, ) { // Each plain-BM25 call returns a bare array of { doc_id, score } const scores = new Map(); for (const h of title) scores.set(h.doc_id, (scores.get(h.doc_id) ?? 0) + wTitle * h.score); for (const h of desc) scores.set(h.doc_id, (scores.get(h.doc_id) ?? 0) + wDesc * h.score); return [...scores.entries()].sort((a, b) => b[1] - a[1]); } for (const [docId, score] of merge(titleHits, descHits).slice(0, 10)) { console.log(docId, score); } ``` #### Go ``` // Index the row's title and description as two separate FTS docs db.FTSIndex(ctx, "shop.products", "title", originchain.FTSIndexRequest{ DocID: "p001", Text: "Wireless Headphones - Pro", }) db.FTSIndex(ctx, "shop.products", "description", originchain.FTSIndexRequest{ DocID: "p001", Text: "Over-ear wireless headphones with noise cancellation.", }) // Query each field independently — each returns a bare []{ DocID, Score } titleHits, _ := db.FTSSearch(ctx, "shop.products", "title", originchain.FTSSearchRequest{ Q: "wireless headphones", Mode: "bm25", K: 20, }) descHits, _ := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "wireless headphones", Mode: "bm25", K: 20, }) // Merge client-side with your own weights const wTitle, wDesc = 2.0, 1.0 scores := map[string]float64{} for _, h := range titleHits { scores[h.DocID] += wTitle * h.Score } for _, h := range descHits { scores[h.DocID] += wDesc * h.Score } ``` what you get back ``` // merged, sorted client-side [ { "doc_id": "p001", "score": 28.92 }, // strong title + description hit { "doc_id": "p027", "score": 14.50 }, // description-only hit { "doc_id": "p014", "score": 9.10 } ] ``` how it works - Each field has its own inverted index and its own document statistics. BM25 on `title` is computed only against other titles - so a match in a 4-word title scores very high. - The two result sets share the `doc_id` namespace because you used the same PK on both POSTs. - You merge by summing weighted scores per `doc_id`. Docs that match in both fields naturally float to the top; docs that match in only one are still present but with their single-field score. common mistakes - Trying to do the merge server-side. There is no multi-field query mode - merging is your app's responsibility. That's deliberate: it keeps the weights tunable without redeploying. - Concatenating fields into one big text blob. You lose the field-level statistics that make title boost work in the first place. Keep them separate. - Using too small a `k`. If a doc only matches one field, it must be inside that field's top-`k` to appear in the merge. Pull a larger `k` per field than you intend to surface. - Using different `doc_id`s across fields. The merge joins on `doc_id`. Use the row's primary key on every field's POST or the merge will silently produce nonsense. --- # Phrase search - OriginChainDB FTS example Canonical source: https://originchaindb.com/docs/examples/fts/phrase Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 4 / 9 # 4. Phrase search [← FTS examples](https://originchaindb.com/docs/examples/fts) what this does Returns documents whose `description` contains the exact phrase `noise cancellation` - both tokens, in that order, with nothing between them. Unlike boolean AND, position matters. when to use it - Branded names and product models: `"Bose QC45"`, `"Sony WH-1000XM5"`. - Multi-word concepts where the order carries the meaning: `"connection refused"`, `"out of memory"`. - Log-template matching - searching for the literal stem of an emitted log line. the request Assumes you have indexed at least a few documents - see [Example 1](https://originchaindb.com/docs/examples/fts/index-doc). #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=noise cancellation" \ --data-urlencode "mode=phrase" ``` #### Python ``` # Phrase search returns a bare JSON array of doc_id strings doc_ids = db.fts.search( "shop.products", "description", q="noise cancellation", mode="phrase", ) for doc_id in doc_ids: print(doc_id) ``` #### TypeScript ``` // Phrase search returns a bare string[] of doc_ids const docIds = await db.ftsSearch("shop.products", "description", { q: "noise cancellation", mode: "phrase", }); for (const docId of docIds) { console.log(docId); } ``` #### Go ``` // Phrase search returns []string of doc_ids docIDs, _ := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "noise cancellation", Mode: "phrase", }) for _, docID := range docIDs { fmt.Println(docID) } ``` what you get back ``` ["p001"] ``` Same shape as boolean: a bare array of `doc_id` strings, no wrapper, no scores. Phrase mode just narrows which docs qualify. how it works - The query is tokenised the same way as the indexed text. - The engine pulls posting lists for each token along with token positions inside each document. - A doc matches only if the positions are consecutive: position(ti+1) = position(ti) + 1 for every adjacent token pair. common mistakes - Expecting stemming to bridge word forms. The API analyzer is Unicode-tokenize + lowercase only - it does not stem. `q=noise cancel` will not match a doc that says "noise cancellation"; the literal token `cancellation` must be present. - Punctuation between tokens. Most analysers strip punctuation - `"noise, cancellation"` in the source text still phrase-matches `q=noise cancellation`. - Expecting fuzzy behaviour. Phrase is strict - one typo and the phrase doesn't match. Use BM25 with `fuzzy=1` if you need typo tolerance. --- # Autocomplete and did-you-mean - OriginChainDB FTS Canonical source: https://originchaindb.com/docs/examples/fts/suggest Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 8 / 9 # 8. Autocomplete and did-you-mean [← FTS examples](https://originchaindb.com/docs/examples/fts) ## what this does One endpoint carries three suggesters, chosen by the `kind` field. `completion` finishes the word a user is still typing, `term` offers a correction for a single misspelled word, and `phrase` rewrites a whole query using the corrections that actually occur together in your corpus. ## when to use it - `completion` - the dropdown under a search box, fired on each keystroke. It reads the term dictionary, never a posting block. - `term` - the "did you mean" line above a thin result set, when one word is clearly wrong. - `phrase` - a multi-word query where every word is plausible on its own but the combination is not. ## the request Autocomplete against an indexed field. Like `_aggs` and `_search`, `_suggest` hangs off the schema and takes the field in the body. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/_suggest" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "kind": "completion", "field": "description", "prefix": "head", "size": 5 }' ``` #### Python ``` import requests # One endpoint, three suggesters. "kind" selects which. resp = requests.post( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/fts/shop.products/_suggest", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "kind": "completion", "field": "description", "prefix": "head", "size": 5, }, ) suggest = resp.json() for c in suggest["completions"]: print(c["text"], c["df"]) ``` #### TypeScript ``` // One endpoint, three suggesters. "kind" selects which. const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/fts/shop.products/_suggest`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ kind: "completion", field: "description", prefix: "head", size: 5, }), }, ); const suggest = await res.json(); for (const c of suggest.completions) { console.log(c.text, c.df); } ``` #### Go ``` // One endpoint, three suggesters. "kind" selects which. body, _ := json.Marshal(map[string]any{ "kind": "completion", "field": "description", "prefix": "head", "size": 5, }) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, "https://"+ocHost+"/v1/tenants/"+ocTenant+"/fts/shop.products/_suggest", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+ocToken) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` { "kind": "completion", "prefix": "head", "completions": [ { "text": "headphones", "df": 3, "total_tf": 4, "score": 3.0 }, { "text": "headset", "df": 1, "total_tf": 1, "score": 1.0 } ], "source": "segmented", "frequency_ranked": true, "df_exact": true, "scanned": 2 } ``` The response echoes the `kind` it was asked for, so all three fit one client type. `df` is the number of documents containing the term and `score` is that same count as a float - the library scores a completion by its document frequency rather than inventing a formula. ## the other two suggesters Same URL, same auth, different `kind`. A term suggestion reports the typed word, its own document frequency, and the candidates within the edit budget: ``` { "kind": "term", "field": "description", "text": "headphnes", "max_edits": 1 } ``` ``` { "kind": "term", "entries": [ { "text": "headphnes", "df": 0, "options": [ { "text": "headphones", "distance": 1, "df": 3, "score": 0.82 } ], "candidates_truncated": false } ], "source": "segmented", "frequency_ranked": true, "df_exact": true } ``` A phrase suggestion corrects every term at once and ranks the whole rewrite, so it can prefer a candidate that is a worse individual match but a far better neighbour: ``` { "kind": "phrase", "field": "description", "text": "wireles headphnes", "max_edits": 1 } ``` ``` { "kind": "phrase", "input": ["wireles", "headphnes"], "input_score": 0.0, "options": [ { "text": "wireless headphones", "terms": ["wireless", "headphones"], "score": 0.61, "term_score": 0.74, "cooccurrence_score": 0.48, "corrected_terms": 2 } ], "source": "segmented", "cooccurrence_available": true, "df_exact": true, "cooccurrence_probes": 4 } ``` ## how it works - The prefix is analyzed, and the last token wins. `"noise cance"` completes `cance`, not the whole string. The analyzed prefix is echoed back so you can see what was actually used. - Defaults. `size` 5, `max_edits` 2, `mode` `missing`, `min_word_length` 4; phrase adds `candidates_per_term` 4 and `confidence` 1.0. - Ordering is stable. Completions sort by document frequency, then total term frequency, then the term itself - the trailing lexicographic key means equal-frequency completions come back in the same sequence every run. `"order": "lexicographic"` turns ranking off when you do your own. - Every response carries honesty flags. `source` says whether frequencies came from the segmented dictionary or a legacy presence set; `frequency_ranked` says whether the list is ranked at all; `df_exact` is true only when no segment holds deletions; `candidates_truncated` warns that a fuzzy expansion was cut before the right answer could be considered. - Caps refuse rather than clamp. A `size` above `MAX_SUGGEST_SIZE` (100) or a `max_scan` above `MAX_PREFIX_SCAN` (100,000) is a 400 - a clamped scan would return a top-k computed over an arbitrary slice of the dictionary. ## common mistakes - Expecting a correction for a correctly-spelled word. The default `mode` is `missing`: a term already in the dictionary is not second-guessed, because a suggester that always has something to say trains people to ignore it. Use `popular` or `always` if you want otherwise. - Treating an empty list as an error. A prefix nothing starts with is 200 with `"completions": []`. The field is indexed, so "no completions" is a true answer - only an unindexed field is a 404. - Trusting `df` as a live count. It counts postings, including postings whose document has since been deleted, so it is an upper bound. `df_exact` tells you when that gap is closed; re-check with a real query if it matters. - Ignoring `frequency_ranked`. When it is false the list is alphabetical rather than best-first, and showing the first entry as "did you mean" is then arbitrary. - Sorting term options by `score` yourself. The list already arrives best-first - nearest edit distance, then frequency. Re-sorting on a blended score usually just reproduces it. --- # Synonyms and stopwords - OriginChainDB FTS example Canonical source: https://originchaindb.com/docs/examples/fts/synonyms-stopwords Sitemap last modified: 2026-09-13T20:18:50.000Z examples · fts · 9 / 9 # 9. Synonyms and stopwords [← FTS examples](https://originchaindb.com/docs/examples/fts) ## what this does Two admin endpoints that configure how one FTS field treats words. A synonym map makes terms interchangeable, so a search for `tv` reaches a document that only ever said `television`. A stopword override replaces the built-in list of words the analyzer discards. Both are keyed by `(schema, field)` and both replace the stored record in full. ## when to use it - Vocabulary no stemmer will ever bridge: `tv` and `television`, a part number and its spelled-out name, an internal codename and the public product name. - Recall complaints after launch. Query-time expansion reaches documents that were indexed long before the map existed, so you can fix a gap without reindexing. - A corpus where the built-in stopword list drops a word that carries meaning in your domain - or keeps one that does not. ## installing a synonym map The body maps each canonical term to the words that should mean the same thing. Send the complete desired map every time - there is no merge. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description/synonyms" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "synonyms": { "television": ["tv", "telly"], "headphones": ["headset", "cans"] } }' ``` #### Python ``` import requests # The COMPLETE map for (schema, field) - this call replaces any previous one. resp = requests.post( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}" f"/fts/shop.products/description/synonyms", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "synonyms": { "television": ["tv", "telly"], "headphones": ["headset", "cans"], } }, ) assert resp.status_code == 201 ``` #### TypeScript ``` // The COMPLETE map for (schema, field) - this call replaces any previous one. const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}` + `/fts/shop.products/description/synonyms`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ synonyms: { television: ["tv", "telly"], headphones: ["headset", "cans"], }, }), }, ); // 201 Created, empty body console.log(res.status); ``` #### Go ``` // The COMPLETE map for (schema, field) - this call replaces any previous one. body, _ := json.Marshal(map[string]any{ "synonyms": map[string][]string{ "television": {"tv", "telly"}, "headphones": {"headset", "cans"}, }, }) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, "https://"+ocHost+"/v1/tenants/"+ocTenant+ "/fts/shop.products/description/synonyms", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+ocToken) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what you get back `201 Created` with an empty body. Both installers answer the same way; a malformed map is a 400. Nothing needs restarting - the records are read per request, so the next query already sees them: ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=tv" \ --data-urlencode "mode=bm25" ``` That search now ranks documents containing `television` as well as those containing `tv`, and the reverse query does the same. ## installing stopwords The array is the complete stopword list for the field. When set, it is used instead of the built-in list for the analyzer language. ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description/stopwords" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "stopwords": ["the", "a", "with", "for"] }' ``` An empty array is not a no-op - it is how you turn stopword removal off for the field entirely: ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description/stopwords" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "stopwords": [] }' ``` ## how it works - Synonyms work from both ends. At index time every occurrence of a canonical term also writes postings under each of its synonyms. At query time a term that appears as either a canonical or a synonym expands to the whole class. That is what makes `tv` and `television` symmetric without writing the index twice. - Expansion is capped at 32 synonyms per class to bound query-time work - a term mapping to a thousand others would otherwise issue a thousand posting-list scans. - Stopwords apply on the analyzer-aware path. They are consulted when a document is indexed with an `analyzer` block, and by the query nodes that carry an analyzer, so that both sides tokenize the same way. An index write with no analyzer uses the default tokenizer and the override does not enter into it. - Both are configuration, not content. They silently change what every later query on the field matches and how it scores, for every reader, without touching a single document - so they are charged as a schema-level change rather than a row write, and need the corresponding privilege. - Each record is read per request, so an install takes effect on the next index or query with no restart and no reindex. ## common mistakes - Treating either call as an append. This is the one that bites. Both replace the stored record in full, so installing one new class silently drops every class you installed before it. Keep the map in source control and post the whole thing. - Expecting an empty stopword array to be ignored. It means "no stopwords for this field", not "leave things as they were" - which is exactly how you disable the built-in list, and exactly how you erase a list you meant to keep. - Expecting a stopword change to rewrite existing postings. It governs analysis from the next write onward. Documents already indexed keep the postings they were given; reindex them if you need the old ones gone. - Installing on the wrong field. Both records are keyed by `(schema, field)`. A map installed on `title` does nothing at all for a query against `description`. - Writing a class longer than the cap. Only the first 32 members of a class survive, so a generated map built from a thesaurus dump will quietly lose its tail. --- # Graph traversal and algorithm examples - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph # Graph examples [← All examples](https://originchaindb.com/docs/examples) OriginChainDB exposes graph traversal and the graph algorithms built on top of it, all over the same declared relations. The traversals - neighbors, reverse, bfs, path and dijkstra - read an adjacency index laid down at write time rather than computed per query. The algorithms above them cover ranked paths, community detection, centrality and node embeddings. 17 work today. Each example below is its own page, with the request, the response shape, side-by-side cURL / Python / TypeScript / Go, and the mistakes that produce a 404 or an empty result. prerequisite Every endpoint below requires a `[[relations]]` block on the schema - the relation must be declared before any traversal can resolve. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations) for the shape. Without a matching relation the call is refused rather than quietly returning an empty result. [1 neighbors - one-hop forward adjacency works today Direct adjacency-list lookup. The fastest graph query on the substrate - one hop, one read. →](https://originchaindb.com/docs/examples/graph/neighbors)[2 reverse - one-hop inbound adjacency works today Who points at this row? The mirror of neighbors. Requires bidirectional=true on the relation (default). →](https://originchaindb.com/docs/examples/graph/reverse)[3 bfs - depth-bounded multi-hop works today Every node reachable within max_depth hops, with the depth at which each was found. →](https://originchaindb.com/docs/examples/graph/bfs)[4 path - reachability check works today Is X connected to Y at all? Returns a boolean. Short-circuits on the first match. →](https://originchaindb.com/docs/examples/graph/path)[5 dijkstra - weighted shortest path works today Single shortest-path cost between two nodes under a caller-supplied edge-weight map. →](https://originchaindb.com/docs/examples/graph/dijkstra)[6 all_simple_paths - every node-disjoint path works today Every distinct path from src to dst. Hard-capped at max_paths to avoid exponential blowup. →](https://originchaindb.com/docs/examples/graph/all-simple-paths)[7 pagerank - over a caller-supplied node set works today PageRank scores over a subgraph you scope explicitly via the nodes= parameter. →](https://originchaindb.com/docs/examples/graph/pagerank)[8 triangles - every closed triple in a relation works today Every closed triple in the relation, walked as undirected. One row per triangle, keys in canonical order. →](https://originchaindb.com/docs/examples/graph/triangles)[9 components - split the graph into islands works today Union-find over the whole node universe. One entry per node carrying the id of the island it landed in. →](https://originchaindb.com/docs/examples/graph/components)[10 louvain - modularity community detection works today The dense groups inside a connected graph, as deterministic integer ids in a communities envelope. →](https://originchaindb.com/docs/examples/graph/louvain)[11 label_propagation - fast seeded communities works today Near-linear community detection. Order-sensitive, so pass a seed when you need to reproduce a run. →](https://originchaindb.com/docs/examples/graph/label-propagation)[12 betweenness - the bridges in a network works today Which nodes sit on the most shortest paths. Exact by default, with an opt-in sampled mode for big graphs. →](https://originchaindb.com/docs/examples/graph/betweenness)[13 eigenvector_centrality - influence by association works today Influence weighted by the influence of your neighbours. No node universe to supply, unlike pagerank. →](https://originchaindb.com/docs/examples/graph/eigenvector-centrality)[14 k-shortest - the top K routes, cheapest first works today Yen's ranked alternatives between two nodes. Note the source and target parameter names, not src and dst. →](https://originchaindb.com/docs/examples/graph/k-shortest)[15 random-walk - seeded and node2vec-biased walks works today A deterministic walk from a start node, with optional p and q bias. The seed is a required parameter. →](https://originchaindb.com/docs/examples/graph/random-walk)[16 node2vec - train embeddings, then query them works today Learn a vector per node from biased walks, persist it, then ask for the nodes most similar to any node. →](https://originchaindb.com/docs/examples/graph/node2vec)[17 graphsage - embeddings that use your columns works today Embeddings that combine the graph with a numeric feature column on each row. Train, persist, then query. →](https://originchaindb.com/docs/examples/graph/graphsage) --- # all_simple_paths - graph path enumeration - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/all-simple-paths Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 6 / 17 # 6. all_simple_paths - every node-disjoint path [← Graph examples](https://originchaindb.com/docs/examples/graph) what this does all_simple_paths enumerates every distinct simple path from `src` to `dst`, up to `max_depth` hops. "Simple" means no node repeats inside a single path. Hard-capped at `max_paths` (default 100) so dense graphs can't blow out the response. when to use it - Citation networks - how does paper A cite paper B, through which intermediates? - Dependency analysis - all routes through which X imports Y. - Exploration / explainability - the surface to show a human "here are all the ways these two are connected". schema requirement The relation must be declared in the schema's `[[relations]]` block. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). the request #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/all_simple_paths" \ --data-urlencode "rel=bought_product" \ --data-urlencode "src=o001" \ --data-urlencode "dst=p001" \ --data-urlencode "max_depth=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` result = db.graph.all_simple_paths( "shop.orders", rel="bought_product", src="o001", dst="p001", max_depth=3, ) ``` #### TypeScript ``` const result = await db.graph.allSimplePaths("shop.orders", { rel: "bought_product", src: "o001", dst: "p001", max_depth: 3, }); ``` #### Go ``` result, _ := db.Graph().AllSimplePaths(ctx, "shop.orders", originchain.AllSimplePathsRequest{ Rel: "bought_product", Src: "o001", Dst: "p001", MaxDepth: 3, }) ``` what you get back ``` [ { "pks": ["o001", "p014", "p001"] }, { "pks": ["o001", "p027", "u091", "p001"] } ] ``` An array of `{ pks: [...] }` objects, one per discovered path. `pks` includes both endpoints, in walk order. how it works - Bounded DFS from `src`, tracking the path stack to enforce the "no repeated node" rule. - Stops descending past `max_depth` and stops emitting past `max_paths`. - The number of simple paths can grow combinatorially with depth - that's why both caps exist. common mistakes - Not capping `max_depth`. The path count is exponential in depth. Even 4-5 hops on a moderately dense graph can hit the `max_paths` cap immediately. - Using this when you only need one. If you just want the shortest, use [dijkstra](https://originchaindb.com/docs/examples/graph/dijkstra). If you just want yes/no, use [path](https://originchaindb.com/docs/examples/graph/path). --- # betweenness - find the bridges in a graph - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/betweenness Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 12 / 17 # 12. betweenness - the bridges in a network [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does betweenness scores each node by how many shortest paths run through it - the centrality that finds bridges and brokers rather than merely popular nodes. `GET /v1/tenants/:t/graph/:schema/betweenness` returns one row per node, sorted by score descending. It is exact by default, and exact betweenness is expensive, so the endpoint refuses on predicted work rather than letting a large graph run past its request budget. ## when to use it - Finding the nodes whose removal would split the network. This is the single most useful centrality for resilience and for fraud work. - Ranking brokers - accounts, services or documents that sit between clusters rather than inside one. - Prioritising review. A bridge with a small degree is much more interesting than a hub with a large one, and only betweenness tells them apart. ## schema requirement The relation must be declared in the schema's `[[relations]]` block and must target the table it is declared on. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request `rel` is required. `max_nodes` caps the node universe and defaults to the engine's 100,000-node ceiling. `samples=K` switches to the Brandes-Pich estimate over K random pivots and `seed` fixes which pivots get chosen. Exact is the default and is never substituted automatically: an approximation you did not ask for and cannot detect would be a correctness bug, so the sampled mode has to be requested explicitly. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/betweenness" \ --data-urlencode "rel=follows" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` scores = db.graph.betweenness("social.users", rel="follows") # The wire `pk` is an array, so the SDK keys the dict by its JSON # encoding: {'["dave"]': 8.0, '["carol"]': 5.0, ...} # # The SDK exposes rel and max_nodes only. Call the endpoint # directly when you need samples= or seed=. ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows" }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/betweenness?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const ranked = await res.json(); ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/betweenness?rel=follows" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var ranked []struct { PK []string `json:"pk"` Betweenness float64 `json:"betweenness"` Estimated bool `json:"estimated"` Pivots int `json:"pivots"` } json.NewDecoder(resp.Body).Decode(&ranked) ``` ## what you get back ``` [ { "pk": ["dave"], "betweenness": 8.0 }, { "pk": ["carol"], "betweenness": 5.0 }, { "pk": ["bob"], "betweenness": 1.0 }, { "pk": ["alice"], "betweenness": 0.0 }, { "pk": ["mallory"], "betweenness": 0.0 } ] ``` ``` [ { "pk": ["dave"], "betweenness": 7.6, "estimated": true, "pivots": 500 }, { "pk": ["carol"], "betweenness": 4.8, "estimated": true, "pivots": 500 } ] ``` One row per node, sorted by score descending, with `pk` as the primary-key array. In exact mode - the first body above - each row has exactly two fields: there is no `estimated` and no `pivots` key at all, so a parser written against the exact shape keeps working. Add `samples=K` and you get the second body: every row is tagged `estimated: true` with the pivot count, which is what lets a consumer holding only the response tell that it is looking at an approximation. ## how it works - Brandes' algorithm runs a single-source shortest-path pass from every node and accumulates dependency scores on the way back. - Cost grows with nodes times edges, so this endpoint refuses on predicted work, not on node count. The refusal is a tagged `graph_work_budget_exceeded` 400 that quotes the node and edge counts, the predicted seconds, and the `samples=` value that would fit - it tells you what to do instead of timing out. - Sampled mode runs the same pass from K randomly chosen pivots instead of from every node, which is what makes a large graph fit inside a request. - `seed` fixes pivot selection and has a fixed default, so a sampled result is reproducible even when you do not pass one. - Sampling only helps where the graph has a distinguishable head. Take the ratio of the top score to the median over the rows you get back: where that ratio is large the sampled ranking tracks the exact one closely, and where it is close to one every node has nearly the same betweenness and no ranking of that graph means much - including the exact one. ## common mistakes - Assuming a hub scores highly. Degree and betweenness answer different questions. A node with many neighbours who all know each other is a poor bridge and scores near zero. - Retrying a work-budget 400. It is not a transient failure. The engine is telling you the exact call is too big and naming the `samples=` value that fits - change the request, do not repeat it. - Comparing a sampled score against an exact one. Estimated scores are only approximately on the same scale. Compare rankings rather than values, and check the `estimated` flag before you do either. - Passing `max_nodes=0` or `samples=0`. Both are refused with a 400. An empty universe and a zero-pivot estimate are not meaningful requests. --- # bfs - depth-bounded graph traversal - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/bfs Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 3 / 17 # 3. bfs - depth-bounded multi-hop [← Graph examples](https://originchaindb.com/docs/examples/graph) what this does bfs returns every distinct node reachable within `max_depth` hops of the source, each tagged with the depth at which it was first discovered. GET /v1/tenants/:t/graph/:schema/bfs takes rel, pk and max_depth, walks forward in concentric rings, and visits each node at most once. when to use it - "Everything within N hops" - friends-of-friends, suggested items, downstream blast radius. - When you need the depth, not just the set - e.g. to weight closer hits higher. - As a pre-filter for a more expensive scoring pass on a smaller candidate set. schema requirement The relation must be declared in the schema's `[[relations]]` block. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). the request #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/bfs?rel=bought_product&pk=o001&max_depth=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` result = db.graph.bfs( "shop.orders", rel="bought_product", pk="o001", max_depth=3, ) ``` #### TypeScript ``` const result = await db.graph.bfs("shop.orders", { rel: "bought_product", pk: "o001", max_depth: 3, }); ``` #### Go ``` result, _ := db.Graph().BFS(ctx, "shop.orders", originchain.BFSRequest{ Rel: "bought_product", PK: "o001", MaxDepth: 3, }) ``` what you get back ``` [ { "pk": "p001", "depth": 1 }, { "pk": "p014", "depth": 1 }, { "pk": "u091", "depth": 2 }, { "pk": "u117", "depth": 2 }, { "pk": "c003", "depth": 3 } ] ``` An array of `{ pk, depth }` objects. Depth is the shortest hop-count from the source; the source itself isn't in the result. how it works - The engine maintains a visited set and a frontier queue, expanding one depth at a time. - Each ring is built by a batched adjacency lookup against the previous frontier - one point-read per node, no scans. - Traversal stops when the frontier is empty or depth reaches `max_depth`. common mistakes - Forgetting `max_depth`. On dense graphs, a missing or very large depth can return millions of nodes. Set a real bound - 2 or 3 is usually enough. - Treating depth as cost. bfs counts hops, not weights. If your edges have meaningful weights, use [dijkstra](https://originchaindb.com/docs/examples/graph/dijkstra). --- # components - find the disconnected islands - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/components Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 9 / 17 # 9. components - split the graph into islands [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does components partitions every row under a schema into connected islands. `GET /v1/tenants/:t/graph/:schema/components` walks `rel` as undirected, runs union-find over the whole node universe, and returns one entry per node carrying the id of the component it landed in. Every row comes back, including rows with no edges at all - those are components of one. ## when to use it - "Is this one network or several?" - the first thing to establish before running any global algorithm over the graph. - Orphan detection. A node alone in its component has no relationships at all, which is usually either a data-quality problem or the answer you were looking for. - Partitioning work. Run an expensive per-island analysis on each component independently instead of on the whole graph at once. ## schema requirement The relation must be declared in the schema's `[[relations]]` block and must target the table it is declared on. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request `rel` names the relation. The optional `include_rows=true` switch also echoes each row's full body alongside its component id; leave it off unless you need the columns, because the default response carries only what the algorithm actually computed. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/components" \ --data-urlencode "rel=follows" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` # The Python SDK does not wrap this endpoint - call it directly. import httpx r = httpx.get( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/graph/social.users/components", params={"rel": "follows"}, headers={"Authorization": f"Bearer {OC_TOKEN}"}, ) for row in r.json(): print(row["pk"][0], "->", row["_component"]) ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows" }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/components?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const islands = new Map(); for (const row of await res.json()) { const key = row._component; islands.set(key, [...(islands.get(key) ?? []), row.pk[0]]); } ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/components?rel=follows" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var rows []struct { PK []string `json:"pk"` Component string `json:"_component"` } json.NewDecoder(resp.Body).Decode(&rows) ``` ## what you get back ``` [ { "pk": ["alice"], "_component": "[\"alice\"]" }, { "pk": ["bob"], "_component": "[\"alice\"]" }, { "pk": ["carol"], "_component": "[\"alice\"]" }, { "pk": ["heidi"], "_component": "[\"heidi\"]" }, { "pk": ["ivan"], "_component": "[\"heidi\"]" }, { "pk": ["judy"], "_component": "[\"heidi\"]" }, { "pk": ["mallory"], "_component": "[\"mallory\"]" } ] ``` One entry per node. `pk` is the node's primary-key array. `_component` is the canonical JSON encoding of the component root's primary key, which makes it a string containing a JSON array, not an array. Treat it as an opaque grouping key: nodes that share it are in the same island. With `include_rows=true` each entry additionally carries the row's own columns, next to the same `_component` field. ## how it works - Union-find over the schema's whole node universe, with every `rel` edge treated as undirected. - Each row is emitted exactly once whether or not it has edges, so the number of distinct `_component` values is the number of islands. - The compact `{pk, _component}` shape is the default because the algorithm's entire answer is one id per node. Echoing row bodies made the response grow with row width without adding information. - `include_rows=true` restores the earlier verbatim-row shape for callers that depended on it. The partitioning is identical either way - only the payload changes. ## common mistakes - Treating `_component` as a key you can look up. It is a canonical-JSON string, not a bare primary key and not an integer. Group on it; do not try to resolve it back to a row. - Reaching for `include_rows=true` by habit. It multiplies the response by your row width. Get the partition first, then fetch only the rows you actually want by primary key. - Expecting isolated rows to be absent. A row with no edges is a component of one and is present in the response. --- # dijkstra - weighted shortest path - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/dijkstra Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 5 / 17 # 5. dijkstra - weighted shortest path [← Graph examples](https://originchaindb.com/docs/examples/graph) what this does dijkstra returns the total cost of the cheapest path from `src` to `dst` under your edge weights, and `null` when the destination is unreachable. GET /v1/tenants/:t/graph/:schema/dijkstra takes rel, src, dst and weights_json. when to use it - Routing / cost minimisation - latency, price, distance. - "Cheapest connection" answers where edges aren't equally costly. - When BFS hop-count isn't the right metric, but you still only need one path. schema requirement The relation must be declared in the schema's `[[relations]]` block. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). the request `weights_json` is a URL-encoded JSON object keyed by `"src|dst"` with a numeric weight. An empty `{}` means "every edge weighs 1" and Dijkstra degenerates to shortest hop-count. #### cURL ``` WEIGHTS='{"o001|p014":1.0,"p014|u091":0.5,"u091|p001":2.0}' curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/dijkstra" \ --data-urlencode "rel=bought_product" \ --data-urlencode "src=o001" \ --data-urlencode "dst=p001" \ --data-urlencode "weights_json=$WEIGHTS" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` weights = { "o001|p014": 1.0, "p014|u091": 0.5, "u091|p001": 2.0, } result = db.graph.dijkstra( "shop.orders", rel="bought_product", src="o001", dst="p001", weights=weights, ) ``` #### TypeScript ``` const weights = { "o001|p014": 1.0, "p014|u091": 0.5, "u091|p001": 2.0, }; const result = await db.graph.dijkstra("shop.orders", { rel: "bought_product", src: "o001", dst: "p001", weights, }); ``` #### Go ``` weights := map[string]float64{ "o001|p014": 1.0, "p014|u091": 0.5, "u091|p001": 2.0, } result, _ := db.Graph().Dijkstra(ctx, "shop.orders", originchain.DijkstraRequest{ Rel: "bought_product", Src: "o001", Dst: "p001", Weights: weights, }) ``` what you get back ``` { "cost": 3.5 } ``` One field: `cost: number | null`. `null` means no path exists. The endpoint returns the cost only - if you need the path itself, use [all_simple_paths](https://originchaindb.com/docs/examples/graph/all-simple-paths). how it works - Classic Dijkstra over a min-heap, expanding the cheapest frontier node. - Edges missing from `weights_json` fall back to a weight of 1. - For the top-K cheapest paths instead of a single shortest, use the [k-shortest endpoint](https://originchaindb.com/docs/graph) - loop-free paths in increasing weight. common mistakes - Negative weights. Dijkstra assumes non-negative edges. Negative weights produce undefined results - reformulate as a positive cost, or pick a different algorithm. - Forgetting to URL-encode `weights_json`. The JSON contains `{}`, quotes, and the `|` separator - use `--data-urlencode` or your SDK's helper rather than concatenating into the URL. --- # eigenvector_centrality - influence scores - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/eigenvector-centrality Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 13 / 17 # 13. eigenvector_centrality - influence by association [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does eigenvector_centrality scores a node by the scores of the nodes attached to it, so being pointed at by three important nodes counts for more than being pointed at by thirty unimportant ones. `GET /v1/tenants/:t/graph/:schema/eigenvector_centrality` runs power iteration over the relation and returns one row per node, sorted descending. Unlike `pagerank`, it needs no node universe: it scores whatever is in the table. ## when to use it - Ranking influence across a whole table without having to choose the candidate set first, which is what `pagerank` makes you do. - Prestige-style scoring, where an endorsement from an already-central node should count for more than one from the periphery. - A cheap global ranking used to seed a more expensive per-candidate model. ## schema requirement The relation must be declared in the schema's `[[relations]]` block and must target the table it is declared on. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request `rel` is required. `max_iter` (default 50) bounds the power iteration and `tol` (default `1e-6`) is the convergence threshold. `max_iter=0` is refused with a 400. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/eigenvector_centrality" \ --data-urlencode "rel=follows" \ --data-urlencode "max_iter=50" \ --data-urlencode "tol=1e-6" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` # The Python SDK does not wrap this endpoint - call it directly. import httpx r = httpx.get( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}" "/graph/social.users/eigenvector_centrality", params={"rel": "follows", "max_iter": 50, "tol": 1e-6}, headers={"Authorization": f"Bearer {OC_TOKEN}"}, ) top = [(row["pk"][0], row["eigenvector"]) for row in r.json()[:10]] ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows", max_iter: "50", tol: "1e-6", }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/eigenvector_centrality?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const ranked = await res.json(); ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/eigenvector_centrality?rel=follows&max_iter=50" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var ranked []struct { PK []string `json:"pk"` Eigenvector float64 `json:"eigenvector"` } json.NewDecoder(resp.Body).Decode(&ranked) ``` ## what you get back ``` [ { "pk": ["dave"], "eigenvector": 0.51 }, { "pk": ["carol"], "eigenvector": 0.44 }, { "pk": ["bob"], "eigenvector": 0.36 }, { "pk": ["alice"], "eigenvector": 0.29 }, { "pk": ["mallory"], "eigenvector": 0.0 } ] ``` One row per node, sorted by score descending. `pk` is the primary-key array and the score field is named `eigenvector`. The score field name differs on every centrality endpoint - `betweenness` returns `betweenness` and `pagerank` returns `score` - so a shared parser has to know which endpoint produced the body it is holding. ## how it works - Power iteration: start from a uniform vector, multiply repeatedly by the adjacency matrix, renormalise, and stop when the change falls below `tol` or after `max_iter` rounds. - The iteration runs on the adjacency matrix plus the identity rather than the adjacency matrix alone. That shift is what stops a bipartite graph oscillating between two states and never converging. - Scores are relative, not absolute. Only the ordering and the ratios between scores carry meaning. - No node universe is required, which is the practical difference from `pagerank`: pagerank makes you name the subgraph up front, this endpoint scores the table as it stands. ## common mistakes - Expecting pagerank's numbers. Eigenvector centrality has no damping factor and no random-restart term, so it concentrates far more weight on the densest region of the graph. The two rankings will differ and neither is wrong. - Reading a zero as a bug. A node with no inbound edges receives no weight. Isolated rows legitimately score zero. - Reusing a betweenness parser. The score field is called `eigenvector` here, and reading `row.betweenness` gives you nothing. --- # graphsage - attribute-aware node embeddings - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/graphsage Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 17 / 17 # 17. graphsage - embeddings that use your columns [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does GraphSAGE learns node vectors from two things at once: the shape of the graph and a numeric feature vector you already store on each row. Where node2vec sees only structure, GraphSAGE aggregates a node's own attributes together with its neighbours'. The shape is the same two calls - `POST /v1/tenants/:t/graph/:schema/graphsage` trains, and `GET /v1/tenants/:t/graph/:schema/graphsage/:rel/topk` queries the persisted result. ## when to use it - Nodes whose columns carry real signal - a price, a score set, an embedding you already computed - where structure alone is not enough. - Cold-start cases. A node with very few edges can still be placed sensibly if its features are informative. - Any recommendation problem where you want structure and content in a single vector instead of blending two separate rankings. ## schema requirement Beyond the `[[relations]]` block, GraphSAGE needs a per-row feature vector: a column holding an array of numbers, the same width on every row. It can be a declared column or an undeclared JSON array field on the row. The width you pass as `feature_dim` must match that column exactly - a mismatch is a 400, never a silent pad or truncate. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## step 1 of 2 - train and persist `rel` and `feature_col` are the two required fields, and `feature_dim` must equal the real width of that column. `num_layers` is capped at 4 and `neighbor_sample_size` at 50. `mean` is the default. All three aggregators run. mean is fully supported; max_pool and lstm are preview - their forward pass is wired, but their own weights stay at their random starting values because the HTTP path does not train them. Only the node embedding table is learned. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/graphsage" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "rel": "follows", "feature_col": "feat", "feature_dim": 4, "embedding_dim": 64, "hidden_dim": 64, "num_layers": 2, "neighbor_sample_size": 10, "aggregator": "mean", "epochs": 5, "seed": 42, "persist": true }' ``` #### Python ``` result = db.graph.graphsage( "social.users", feature_col="feat", rel="follows", config={ "feature_dim": 4, "embedding_dim": 64, "hidden_dim": 64, "num_layers": 2, "neighbor_sample_size": 10, "aggregator": "mean", "epochs": 5, "seed": 42, }, persist=True, ) ``` #### TypeScript ``` const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/graphsage`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ rel: "follows", feature_col: "feat", feature_dim: 4, embedding_dim: 64, hidden_dim: 64, num_layers: 2, neighbor_sample_size: 10, aggregator: "mean", epochs: 5, seed: 42, persist: true, }), }, ); const trained = await res.json(); ``` #### Go ``` body, _ := json.Marshal(map[string]any{ "rel": "follows", "feature_col": "feat", "feature_dim": 4, "embedding_dim": 64, "hidden_dim": 64, "num_layers": 2, "neighbor_sample_size": 10, "aggregator": "mean", "epochs": 5, "seed": 42, "persist": true, }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+OC_HOST+"/v1/tenants/"+OC_TENANT+"/graph/social.users/graphsage", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` ## what training returns ``` { "embeddings": { "alice": [0.0311, -0.0142, 0.0207], "bob": [0.0288, -0.0166, 0.0193] }, "vocab_size": 11, "training_pairs": 3960, "final_loss": 0.5842, "embedding_dim": 64, "feature_dim": 4, "persisted": true } ``` The same envelope node2vec returns, with two differences: the width field is called `embedding_dim` rather than `dim`, and `feature_dim` echoes the input width back so you can confirm the engine read the column you meant. Vectors are shortened above - each one is `embedding_dim` floats long. ## step 2 of 2 - ask for similar nodes Identical grammar to the node2vec query: the relation is a path segment, `query` names the node and `k` the number of neighbours, and `metric` accepts `cosine`, `dot`, `l2` or `manhattan`, defaulting to cosine. The two persisted sets are stored separately, so a relation can carry both a node2vec and a GraphSAGE index at once. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/graphsage/follows/topk" \ --data-urlencode "query=alice" \ --data-urlencode "k=5" \ --data-urlencode "metric=cosine" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` hits = db.graph.graphsage_topk( "social.users", rel="follows", query_pk="alice", k=5, metric="cosine", ) for hit in hits: print(hit.pk, hit.score) ``` #### TypeScript ``` const qs = new URLSearchParams({ query: "alice", k: "5", metric: "cosine", }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/graphsage/follows/topk?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const hits = await res.json(); ``` #### Go ``` url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/graphsage/follows/topk?query=alice&k=5&metric=cosine" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var hits []struct { PK string `json:"pk"` Score float32 `json:"score"` } json.NewDecoder(resp.Body).Decode(&hits) ``` ## what the query returns ``` [ { "pk": "bob", "score": 0.97 }, { "pk": "dave", "score": 0.88 }, { "pk": "carol", "score": 0.81 } ] ``` A flat array of `{ pk, score }`, best first, with the query node excluded from its own results. Until a training call has run with `persist: true` this route answers 503 with a message pointing at the persistence step - a missing prerequisite rather than a failure. ## how it works - Each layer samples up to `neighbor_sample_size` neighbours and averages their representations with the node's own. That averaging is the `mean` aggregator, and `num_layers` decides how far out the neighbourhood reaches. - Features are read per row from the feature column. Both a declared column and an undeclared JSON array field on the row work. - The same seed, graph, features and configuration give byte-identical vectors. - `persist: true` writes one blob per schema and relation, stored alongside but separately from the node2vec blob, so both indexes can exist for the same relation. - Response size is capped exactly as node2vec is: `vocab_size` times `embedding_dim` above one million floats is a 400. - A persisted set that has gone bad can be inspected with a `health` call on the same path and retrained in place with a `rebuild` call, which retrains from the rows and edges already in the store. No re-ingest is needed, and only the derived embedding blob is rewritten - no row is touched. ## common mistakes - Getting `feature_dim` wrong. Declaring 128 against a four-wide column fails loudly with a 400. That is deliberate: a silently padded feature vector would produce embeddings that look fine and mean nothing. - Pointing `feature_col` at a column that is not there. An unknown name is a 400, not an empty feature vector. - Asking for an aggregator other than `mean`. Assuming max_pool and lstm fail. They do not - they return embeddings. What they do not do is train their own aggregator weights over HTTP, so treat them as preview rather than as a drop-in for mean. - Forgetting `persist: true`. As with node2vec, the topk route answers 503 until a persisted blob exists. - Exceeding `num_layers` 4 or `neighbor_sample_size` 50. Both are refused with a 400 rather than clamped. --- # k-shortest - the top K routes between nodes - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/k-shortest Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 14 / 17 # 14. k-shortest - the top K routes, cheapest first [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does k-shortest returns up to `k` loop-free routes between two nodes, cheapest first, with the node sequence and the cost of each. Where `dijkstra` gives you a single number, this gives you the routes themselves and their runners-up. `GET /v1/tenants/:t/graph/:schema/k-shortest` takes `rel`, `source`, `target` and `k`. ## when to use it - Presenting alternatives - the best route plus the next few, so a person can choose. - Failover planning. If the cheapest path breaks, what is the second cheapest, and does it depend on the same intermediate node? - Explaining a connection. "Here are the three ways A reaches B" is a much better answer than a single cost. ## schema requirement The relation must be declared in the schema's `[[relations]]` block. If you pass `weight_col`, the relation's target table must also declare that column. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request This endpoint names its endpoints `source` and `target`, not `src` and `dst` the way `path`, `dijkstra` and `all_simple_paths` do. `k` is required and is capped at 50; asking for more is a 400 rather than a silent truncation. By default every edge weighs 1, so the ranking is by hop count - pass `weight_col` to read each edge's weight from a named column on the destination row instead. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/k-shortest" \ --data-urlencode "rel=follows" \ --data-urlencode "source=alice" \ --data-urlencode "target=erin" \ --data-urlencode "k=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` paths = db.graph.k_shortest( "social.users", src="alice", # sent as the `source` query parameter target="erin", rel="follows", k=3, ) for p in paths: print(p.nodes, p.cost) ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows", source: "alice", target: "erin", k: "3", }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/k-shortest?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const { paths } = await res.json(); ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/k-shortest?rel=follows&source=alice&target=erin&k=3" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var out struct { Paths []struct { Nodes []string `json:"nodes"` Cost float64 `json:"cost"` } `json:"paths"` } json.NewDecoder(resp.Body).Decode(&out) ``` ## what you get back ``` { "paths": [ { "nodes": ["alice", "carol", "dave", "erin"], "cost": 3.0 }, { "nodes": ["alice", "bob", "carol", "dave", "erin"], "cost": 4.0 } ] } ``` An object with a `paths` array, one entry per route, ascending by `cost`. `nodes` is the walk in order as plain strings, including both endpoints. Fewer than `k` entries simply means the graph has no more loop-free routes. An unreachable target is a 200 with `"paths": []`, not a 404, and a `source` equal to `target` comes back as a single zero-cost path. ## how it works - Yen's algorithm finds the shortest path, then repeatedly forces a deviation at each node along it and keeps the cheapest result it has not already emitted. - Routes are loop-free by construction: no node repeats inside a single path. - With no `weight_col` every edge weighs 1, so the ranking is by hop count and matches what a breadth-first search would find. - With `weight_col` set, each edge's weight is read from that column on the destination row and coerced to a float. The relation has to exist on the schema for that lookup, and an unknown relation name is a 400. - `k` above 50 is rejected with the ceiling quoted in the message, so callers learn to budget instead of silently receiving a truncated list. ## common mistakes - Sending `src` and `dst`. They are `source` and `target` on this endpoint. The names differ from the other path endpoints, and the wrong name is a rejected request. - Treating an empty `paths` array as an error. Unreachable is a normal 200. Check the array length, not the status code. - Expecting `cost` to mean hops once `weight_col` is set. With weights coming from a column, `cost` is the sum of those column values and carries whatever units that column has. - Asking for a large `k` just in case. Above 50 the whole request is refused, so you get nothing rather than the first fifty. --- # label_propagation - fast seeded communities - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/label-propagation Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 11 / 17 # 11. label_propagation - fast seeded communities [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does label_propagation gives every node the label held by most of its neighbours and repeats until the labels stop moving. `GET /v1/tenants/:t/graph/:schema/label_propagation` returns one row per node with the label it converged to. It is far cheaper than louvain and far less stable: the algorithm is order-sensitive, so the engine shuffles visit order from a seeded generator, and the seed is what makes a run repeatable. ## when to use it - A cheap first look at community structure on a graph too large to be worth a modularity pass. - Recomputing a grouping often, where near-linear cost matters more than the quality of the partition. - A cross-check against louvain. Where the two agree the grouping is robust; where they disagree the boundary is genuinely fuzzy. ## schema requirement The relation must be declared in the schema's `[[relations]]` block and must target the table it is declared on. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request `rel` is required and `max_iter` defaults to 20. `seed` is optional in the protocol but effectively mandatory in practice: omit it and the engine picks the current unix timestamp, and because the response does not echo the seed back, an unseeded run can never be reproduced or explained afterwards. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/label_propagation" \ --data-urlencode "rel=follows" \ --data-urlencode "max_iter=20" \ --data-urlencode "seed=42" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` labels = db.graph.label_propagation( "social.users", rel="follows", seed=42, max_iter=20, ) # The wire `pk` is an array, so the SDK keys the dict by its JSON # encoding: {'["alice"]': 0, '["heidi"]': 7, ...} ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows", max_iter: "20", seed: "42", }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/label_propagation?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const rows = await res.json(); ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/label_propagation?rel=follows&max_iter=20&seed=42" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var rows []struct { PK []string `json:"pk"` Label uint64 `json:"label"` } json.NewDecoder(resp.Body).Decode(&rows) ``` ## what you get back ``` [ { "pk": ["alice"], "label": 0 }, { "pk": ["bob"], "label": 0 }, { "pk": ["carol"], "label": 0 }, { "pk": ["dave"], "label": 0 }, { "pk": ["heidi"], "label": 7 }, { "pk": ["ivan"], "label": 7 }, { "pk": ["judy"], "label": 7 }, { "pk": ["mallory"], "label": 10 } ] ``` One row per node, ordered by primary key. `pk` is the primary-key array and `label` is an unsigned integer. Labels start as one-per-node identifiers and most of them die off during convergence, so the surviving values are not contiguous and carry no ordering - the only thing a label means is equality. Two nodes with the same label are in the same community. ## how it works - Every node starts with its own label. Each round visits nodes in a shuffled order and moves each node to whichever label most of its neighbours hold. - The shuffle is the only thing the seed drives. Ties within a tally are broken deterministically by the smaller label, so the same seed on the same graph gives byte-identical output. - The pass repeats until no label changes or `max_iter` rounds have run. There is no convergence guarantee, which is exactly why the iteration cap exists. - Self-loops are ignored - a node never votes for itself by force. - `max_iter=0` is refused with a 400 rather than returning an unconverged answer. ## common mistakes - Omitting `seed`. The engine falls back to the current unix timestamp and does not tell you which one it used. Two calls a second apart can return different groupings with nothing in the response to explain the difference. - Reading labels as group numbers. They are not dense, not ordered and not stable across graphs. Compare them for equality and nothing else. - Expecting louvain's answer. Label propagation optimises nothing globally. On a graph without sharp community structure it can collapse most of the nodes into one label, which is a real result and not an error. Cross-check with louvain before trusting a partition. --- # louvain - modularity community detection - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/louvain Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 10 / 17 # 10. louvain - modularity community detection [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does louvain runs two-phase modularity-greedy community detection over one table's relation graph. `GET /v1/tenants/:t/graph/:schema/louvain` returns a `communities` envelope with one entry per row: the node's primary key and the integer id of the community it settled into. Where `components` only separates nodes that cannot reach each other at all, louvain finds the densely connected groups inside a single connected island. ## when to use it - Segmenting an interaction graph into groups that genuinely talk to each other, rather than groups that merely happen to be reachable. - Topic clustering over a citation or co-occurrence graph, where the grouping should come from the edges rather than from a label you already assigned. - A first pass before an expensive per-group model: run louvain, then treat each community as an independent unit of work. ## schema requirement The relation must be declared in the schema's `[[relations]]` block and must target the table it is declared on. louvain refuses a relation that points at another table, and the error names both tables. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request `rel` is the only required parameter. `tolerance` (default `1e-4`) is the modularity delta below which the outer aggregation loop stops, and `max_levels` (default `10`) caps that loop. Real graphs converge in a handful of levels - the cap is a guard against pathological input, not a knob you normally touch. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/louvain" \ --data-urlencode "rel=follows" \ --data-urlencode "tolerance=1e-4" \ --data-urlencode "max_levels=10" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` communities = db.graph.louvain( "social.users", rel="follows", tolerance=1e-4, max_levels=10, ) # The SDK flattens the envelope to {pk: community_id}: # {"alice": 0, "bob": 0, "carol": 0, "heidi": 1, ...} ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows", max_levels: "10" }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/louvain?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const { communities } = await res.json(); ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/louvain?rel=follows&max_levels=10" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var out struct { Communities []struct { PK string `json:"pk"` Community uint32 `json:"community"` } `json:"communities"` } json.NewDecoder(resp.Body).Decode(&out) ``` ## what you get back ``` { "communities": [ { "pk": "alice", "community": 0 }, { "pk": "bob", "community": 0 }, { "pk": "carol", "community": 0 }, { "pk": "dave", "community": 0 }, { "pk": "heidi", "community": 1 }, { "pk": "ivan", "community": 1 }, { "pk": "judy", "community": 1 }, { "pk": "mallory", "community": 2 } ] } ``` An object with a single `communities` array - not a bare array, and not the row shape the other algorithm endpoints use. Here `pk` is a plain string and `community` is a dense integer in the range zero to k. Ids are assigned in order of first appearance by lexicographic primary key, so the lexicographically smallest key is always in community 0 and a rerun over an unchanged graph produces the same numbering. ## how it works - Phase one moves each node into whichever neighbouring community gives the largest modularity gain. Phase two collapses each community into a single node and repeats. - The loop stops when the modularity delta between levels falls below `tolerance`, when phase one stops moving nodes, or when `max_levels` is reached. - Community ids are dense and deterministic, so two runs over the same graph are directly comparable and can be stored as-is. - The endpoint refuses graphs above its 500,000-node ceiling. The per-level pass is linear in nodes plus edges, and a multi-million-node call will not land inside a request budget. ## common mistakes - Reading the response as an array. It is `{ "communities": [ ... ] }`. Indexing the body directly gets you nothing. - Expecting `pk` to be an array here. louvain returns a plain string, while `components`, `betweenness`, `label_propagation` and `eigenvector_centrality` all return a primary-key array. The response grammar genuinely differs per endpoint - write the parser per endpoint. - Comparing community ids across graphs. The numbering comes from this graph's own key order. Ids from a different graph, or from the same graph after new rows land, are not comparable. --- # neighbors - one-hop graph traversal - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/neighbors Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 1 / 17 # 1. neighbors - one-hop forward adjacency [← Graph examples](https://originchaindb.com/docs/examples/graph) what this does neighbors walks one hop forward along a declared relation: given the primary key of a source row, it returns the array of primary keys that row points at. GET /v1/tenants/:t/graph/:schema/neighbors takes the relation name in rel and the source key in pk. when to use it - "What products did this order buy?" - one hop, forward direction. - Anywhere you would write `SELECT * FROM ... WHERE fk = ?` just to follow a foreign key. - As the inner loop of your own traversal when you want hop-by-hop control on the client. schema requirement The schema for `shop.orders` must declare a `[[relations]]` block with `name = "bought_product"`. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). Without it, the engine returns 404. the request #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/neighbors?rel=bought_product&pk=o001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` result = db.graph.neighbors( "shop.orders", rel="bought_product", pk="o001", ) ``` #### TypeScript ``` const result = await db.graph.neighbors("shop.orders", { rel: "bought_product", pk: "o001", }); ``` #### Go ``` result, _ := db.Graph().Neighbors(ctx, "shop.orders", originchain.NeighborsRequest{ Rel: "bought_product", PK: "o001", }) ``` what you get back ``` [ "p001", "p014", "p027" ] ``` A flat array of target primary keys. Order is the insertion order of the edges; do not rely on it being sorted. how it works - Adjacency is laid down at write time, keyed by `(tenant, relation, source_pk)`. - The read is a single point lookup against that key - no scan, no join. - Cost is independent of total table size; it scales only with the fan-out of this one row. common mistakes - Relation not declared. Adding a foreign-key column isn't enough - the relation must be in the schema's `[[relations]]` table. - Wrong direction. neighbors is forward only. For "who points at this row?", use [reverse](https://originchaindb.com/docs/examples/graph/reverse). --- # node2vec - train node embeddings and query - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/node2vec Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 16 / 17 # 16. node2vec - train embeddings, then query them [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does node2vec learns a vector for every node from biased random walks, so nodes sitting in similar parts of the graph end up close together. It is two calls: `POST /v1/tenants/:t/graph/:schema/node2vec` trains, and `GET /v1/tenants/:t/graph/:schema/node2vec/:rel/topk` answers similarity queries. The second only works if the first was made with `persist: true` - without it the vectors come back in the response and nothing is stored. ## when to use it - "Nodes like this one", where similarity should come from graph position rather than from text or image content. - Link prediction and recommendation over a relation you already model as a foreign key. - Feeding graph structure into a downstream model as a fixed-width vector, without writing a walk sampler yourself. ## schema requirement The relation must be declared in the schema's `[[relations]]` block. node2vec needs no feature columns - it learns from structure alone. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## step 1 of 2 - train and persist `rel` is the only required field; every other knob has a default, so a body of just the relation name is a complete request. Set `persist: true` to store the vectors under the schema and relation so the topk route becomes available. Without it the call is a pure computation: it trains in memory, returns the vectors and writes nothing. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/node2vec" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "rel": "follows", "dim": 64, "walks_per_node": 10, "walk_length": 40, "window_size": 5, "epochs": 5, "seed": 42, "persist": true }' ``` #### Python ``` # The Python SDK wraps the query side but not training - # post the config directly. import httpx r = httpx.post( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/graph/social.users/node2vec", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "rel": "follows", "dim": 64, "walks_per_node": 10, "walk_length": 40, "window_size": 5, "epochs": 5, "seed": 42, "persist": True, }, ) trained = r.json() print(trained["vocab_size"], trained["persisted"]) ``` #### TypeScript ``` const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/node2vec`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ rel: "follows", dim: 64, walks_per_node: 10, walk_length: 40, window_size: 5, epochs: 5, seed: 42, persist: true, }), }, ); const trained = await res.json(); ``` #### Go ``` body, _ := json.Marshal(map[string]any{ "rel": "follows", "dim": 64, "walks_per_node": 10, "walk_length": 40, "window_size": 5, "epochs": 5, "seed": 42, "persist": true, }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+OC_HOST+"/v1/tenants/"+OC_TENANT+"/graph/social.users/node2vec", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` ## what training returns ``` { "embeddings": { "alice": [0.0142, -0.0318, 0.0091], "bob": [-0.0077, 0.0264, -0.0155], "carol": [0.0203, -0.0119, 0.0288] }, "vocab_size": 11, "training_pairs": 4820, "final_loss": 0.6931, "dim": 64, "persisted": true } ``` `embeddings` maps each node's primary key to its vector, and each vector is `dim` floats long - they are shortened above to keep the example readable. `vocab_size` is how many nodes appeared in a walk, `training_pairs` how many skip-gram pairs were consumed, and `final_loss` the loss at the end of the last epoch. `persisted` mirrors the request flag and is true only when the blob was actually written, so you can tell whether the next call will work without probing for it. ## step 2 of 2 - ask for similar nodes The relation name moves into the path here, because that is how the persisted blob is keyed. `query` is the node you are asking about and `k` is how many neighbours you want. `metric` accepts `cosine`, `dot`, `l2` or `manhattan` and defaults to cosine. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/node2vec/follows/topk" \ --data-urlencode "query=alice" \ --data-urlencode "k=5" \ --data-urlencode "metric=cosine" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` hits = db.graph.node2vec_topk( "social.users", rel="follows", query_pk="alice", k=5, metric="cosine", ) ``` #### TypeScript ``` const qs = new URLSearchParams({ query: "alice", k: "5", metric: "cosine", }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/node2vec/follows/topk?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const hits = await res.json(); ``` #### Go ``` url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/node2vec/follows/topk?query=alice&k=5&metric=cosine" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var hits []struct { PK string `json:"pk"` Score float32 `json:"score"` } json.NewDecoder(resp.Body).Decode(&hits) ``` ## what the query returns ``` [ { "pk": "bob", "score": 0.94 }, { "pk": "carol", "score": 0.91 }, { "pk": "dave", "score": 0.72 } ] ``` A flat array of `{ pk, score }`, best first, with the query node itself excluded. `pk` is a plain string here, not the primary-key array the algorithm endpoints return. Fewer than `k` hits means the persisted set is smaller than `k`; asking for more than the engine's node ceiling truncates the result rather than failing the call. ## how it works - Training samples biased random walks from every node, then runs skip-gram with negative sampling over the resulting sequences. `p` and `q` bias the walks exactly as they do on the `random-walk` endpoint. - The same seed, graph and configuration produce byte-identical vectors, so a retrain is reproducible. - `persist: true` writes one blob per schema and relation, and the next persist replaces it atomically. What is stored is the trained vectors themselves rather than the recipe - replaying training elsewhere would not land on the same numbers. - The response is capped at one million floats, which is `vocab_size` times `dim`. Above that the call is refused with a 400, because a body that large stops being something a normal gateway will carry. - `dim` above 1024 is refused, and so is a walk budget above ten million, where the budget is `walks_per_node` multiplied by `walk_length`. ## common mistakes - Forgetting `persist: true`. Training without it returns the vectors and stores nothing, and the topk route then answers 503 telling you to post with persist first. That 503 is a missing prerequisite, not an outage. - Putting the relation in the query string on the topk call. It belongs in the path, as `.../node2vec//topk`. - Misspelling a knob. Unknown request fields are dropped rather than rejected by default, so `walk_len` instead of `walk_length` trains at the default and still returns 200. Check the names against the body above. - Comparing scores across metrics. A cosine similarity and an l2 distance are not on the same scale. Pick one metric and stay with it. --- # pagerank - influence scores on a subgraph - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/pagerank Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 7 / 17 # 7. pagerank - over a caller-supplied node set [← Graph examples](https://originchaindb.com/docs/examples/graph) what this does pagerank computes PageRank scores over a subgraph that you define via `nodes=`. The engine restricts the iteration to that node universe and the edges between them; everything outside is ignored. Returns an array of `{ pk, score }` entries, sorted by score descending and summing to 1. when to use it - Influence ranking inside a tenant / project / topic - "who matters in this subgraph?" - Recommendation seeds - score a candidate set, take the top-K. - Any case where the full graph is too big to rank globally, but a slice is the right scope. schema requirement The relation must be declared in the schema's `[[relations]]` block. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). the request `nodes` is a comma-separated list of the primary keys that form the universe. Edges to keys outside the list are dropped; the iteration is scoped to what you pass. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/pagerank" \ --data-urlencode "rel=bought_product" \ --data-urlencode "nodes=o001,o002,o003,o004,o005" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` result = db.graph.pagerank( "shop.orders", rel="bought_product", nodes=["o001", "o002", "o003", "o004", "o005"], ) ``` #### TypeScript ``` const result = await db.graph.pagerank("shop.orders", { rel: "bought_product", nodes: ["o001", "o002", "o003", "o004", "o005"], }); ``` #### Go ``` result, _ := db.Graph().PageRank(ctx, "shop.orders", originchain.PageRankRequest{ Rel: "bought_product", Nodes: []string{"o001", "o002", "o003", "o004", "o005"}, }) ``` what you get back ``` [ { "pk": "o004", "score": 0.2417 }, { "pk": "o001", "score": 0.2185 }, { "pk": "o003", "score": 0.1962 }, { "pk": "o002", "score": 0.1806 }, { "pk": "o005", "score": 0.1630 } ] ``` An array of { pk, score } entries, already sorted by score descending, so the first entry is the top influencer in the universe you passed. Scores sum to 1 across that node set. how it works - Power-iteration PageRank with damping factor 0.85. - The induced subgraph is built once from the `nodes` list; iteration converges on that subgraph only. - Caller-scoped universe is intentional - the engine refuses to rank the entire tenant graph because that's almost never what you want and is unbounded in cost. common mistakes - Omitting `nodes=`. The engine 400s because it requires the universe up front. Pick the candidate set first; rank second. - Comparing scores across different node sets. Scores are relative to the universe you pass. A pk's score in one subgraph isn't comparable to its score in another. --- # path - graph reachability check - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/path Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 4 / 17 # 4. path - reachability check [← Graph examples](https://originchaindb.com/docs/examples/graph) what this does path returns a single boolean: whether `src` reaches `dst` within `max_depth` hops along a declared relation. The traversal short-circuits the moment it finds the first connecting route - it does not enumerate every path. when to use it - Permission checks - "is this user in the org tree under that admin?" - Reachability gates before running an expensive operation - cheaper than full BFS. - "Are these two entities related?" boolean badges in UIs. schema requirement The relation must be declared in the schema's `[[relations]]` block. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). the request #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/path" \ --data-urlencode "rel=bought_product" \ --data-urlencode "src=o001" \ --data-urlencode "dst=p001" \ --data-urlencode "max_depth=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` result = db.graph.path( "shop.orders", rel="bought_product", src="o001", dst="p001", max_depth=3, ) ``` #### TypeScript ``` const result = await db.graph.path("shop.orders", { rel: "bought_product", src: "o001", dst: "p001", max_depth: 3, }); ``` #### Go ``` result, _ := db.Graph().Path(ctx, "shop.orders", originchain.PathRequest{ Rel: "bought_product", Src: "o001", Dst: "p001", MaxDepth: 3, }) ``` what you get back ``` { "reachable": true } ``` One field: `reachable: true | false`. `false` means "not within `max_depth`" - it isn't proof there's no path at higher depth. how it works - Same BFS frontier as [bfs](https://originchaindb.com/docs/examples/graph/bfs), but it stops the instant `dst` appears in any expansion. - Worst case (no path) costs the same as a bounded bfs; best case is a single hop. - Use this over bfs whenever you only need the boolean - the savings come from not materialising the visited set. common mistakes - Reading `false` as "no path". It only means "no path within `max_depth`". If you need a definitive answer, raise the depth or use [all_simple_paths](https://originchaindb.com/docs/examples/graph/all-simple-paths). - Swapping `src` and `dst` on a one-way relation. path is direction-sensitive. If you need the inverse direction, check that the relation is bidirectional. --- # random-walk - seeded walks over a relation - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/random-walk Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 15 / 17 # 15. random-walk - seeded and node2vec-biased walks [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does random-walk takes a starting node and walks the relation forward for up to `steps` hops, returning the sequence of primary keys it visited. The walk is deterministic: the same seed on the same graph produces the same walk, byte for byte. Passing `p` or `q` switches to the node2vec-biased walk, which is the same sampler the embedding endpoints use internally. ## when to use it - Sampling a large graph cheaply. A few thousand walks give a usable picture without touching every node. - Generating training sequences for your own embedding or sequence model, outside the engine's built-in node2vec. - Reproducible exploration. A seeded walk is something you can paste into a bug report and have someone else replay exactly. ## schema requirement The relation must be declared in the schema's `[[relations]]` block. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request `rel`, `start`, `steps` and `seed` are all required - there is no default seed, because a walk you cannot reproduce is not much use. `steps` is capped at 1000. `p` (the return parameter) and `q` (the in-out parameter) are the node2vec bias knobs; setting either one routes the request through the biased walk, and `p=1` with `q=1` is indistinguishable from the unbiased walk. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/random-walk" \ --data-urlencode "rel=follows" \ --data-urlencode "start=alice" \ --data-urlencode "steps=5" \ --data-urlencode "seed=42" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` walk = db.graph.random_walk( "social.users", start="alice", rel="follows", steps=5, seed=42, ) # The SDK returns just the walk list: # ["alice", "carol", "dave", "erin", "dave", "erin"] ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows", start: "alice", steps: "5", seed: "42", }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/random-walk?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const { start, walk } = await res.json(); ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/random-walk?rel=follows&start=alice&steps=5&seed=42" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var out struct { Start string `json:"start"` Walk []string `json:"walk"` } json.NewDecoder(resp.Body).Decode(&out) ``` ## what you get back ``` { "start": "alice", "walk": ["alice", "carol", "dave", "erin", "dave", "erin"] } ``` `walk` begins with the start node, so its length is at most `steps` plus one. It can be shorter: the walk stops as soon as it reaches a node with no outbound edges. A dead-end start returns a walk containing only that node, and a node with a self-loop walks in place. ## how it works - At each hop the walker picks uniformly among the current node's outbound edges on `rel`, driven by a seeded generator. - The same graph, start, step count and seed always produce the same body, so two identical calls are byte-identical on the wire. - With `p` or `q` present the walk becomes second-order: `p` controls how likely it is to step straight back where it came from, and `q` how likely it is to move away from the previous node's neighbourhood. Setting both to 1 reduces to the uniform walk by construction. - `steps` above 1000 is refused with a 400 quoting the cap, and `p=0` is refused because a zero return parameter is not a valid bias. ## common mistakes - Omitting `seed`. It is a required parameter, not an optional one, and the request is rejected without it. - Assuming the walk has `steps` plus one entries. It stops early at a dead end. Read the length rather than trusting it. - Reading a repeated node as a bug. A walk is not a path. It revisits nodes freely, and a two-cycle will bounce between the same pair. - Drawing conclusions from one walk. A single seeded walk is one draw from a distribution. Aggregate many walks across many seeds before you believe a pattern. --- # reverse - inbound graph adjacency - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/reverse Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 2 / 17 # 2. reverse - one-hop inbound adjacency [← Graph examples](https://originchaindb.com/docs/examples/graph) what this does reverse returns the primary keys of every row that has an edge pointing at a given row - the mirror of [neighbors](https://originchaindb.com/docs/examples/graph/neighbors), same relation, opposite direction. GET /v1/tenants/:t/graph/:schema/reverse takes the relation name in rel and the target key in pk. when to use it - "Which orders bought this product?" - the inverse of the forward query. - Anywhere you would scan a table to find rows whose foreign key matches a value. - Building inbound feeds, audit trails, "referenced by" panels. schema requirement The relation must be declared with `bidirectional = true` - which is the default. If a relation is one-way only, reverse on it returns 400. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). the request #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/reverse?rel=bought_product&pk=p001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` result = db.graph.reverse( "shop.orders", rel="bought_product", pk="p001", ) ``` #### TypeScript ``` const result = await db.graph.reverse("shop.orders", { rel: "bought_product", pk: "p001", }); ``` #### Go ``` result, _ := db.Graph().Reverse(ctx, "shop.orders", originchain.ReverseRequest{ Rel: "bought_product", PK: "p001", }) ``` what you get back ``` [ "o001", "o042", "o118" ] ``` A flat array of source primary keys. Same shape as neighbors, just walking the other direction. how it works - When the relation is bidirectional, the write path lays down a second adjacency entry keyed by the target. - reverse is then a point lookup against that inbound entry - identical cost profile to neighbors. - The cost is the price of bidirectionality: storage roughly doubles for the edge, but reads stay O(fan-in). common mistakes - Relation set to `bidirectional = false`. reverse refuses to run. Re-declare the relation or use forward traversal from the other side. - High fan-in nodes. A "popular" product can have hundreds of thousands of inbound edges. Paginate at the call site or use [pagerank](https://originchaindb.com/docs/examples/graph/pagerank) when you want ranked influence rather than the raw list. --- # triangles - find every closed triple - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/graph/triangles Sitemap last modified: 2026-09-13T20:18:50.000Z examples · graph · 8 / 17 # 8. triangles - every closed triple in a relation [← Graph examples](https://originchaindb.com/docs/examples/graph) ## what this does triangles enumerates every closed triple in one table's own relation graph. `GET /v1/tenants/:t/graph/:schema/triangles` takes a single `rel` parameter, symmetrises the relation so an edge counts in both directions, and returns one row per triangle with the three primary keys in canonical order. ## when to use it - Mutual-connection detection. Three accounts that all reference each other is a far stronger signal than three that merely share a neighbour. - Local density analytics. Triangle count over a node's degree is the standard clustering-coefficient measure. - Ring detection in payments, referrals and reciprocal-link networks, where the closed triple is exactly the shape you are hunting for. ## schema requirement The relation must be declared in the schema's `[[relations]]` block and must target the table it is declared on - triangle enumeration walks one table's own edge set, and a relation pointing at another table is refused. See [schemas/reference#relations](https://originchaindb.com/docs/schemas/reference#relations). ## the request `rel` is the only parameter this endpoint accepts. There is no depth, no start node and no cap: the answer is every triangle in the relation. #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.users/triangles" \ --data-urlencode "rel=follows" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` # The Python SDK does not wrap this endpoint - call it directly. import httpx r = httpx.get( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/graph/social.users/triangles", params={"rel": "follows"}, headers={"Authorization": f"Bearer {OC_TOKEN}"}, ) triangles = r.json() ``` #### TypeScript ``` // The TypeScript SDK does not wrap the graph algorithms - use fetch. const qs = new URLSearchParams({ rel: "follows" }); const res = await fetch( `https://${OC_HOST}/v1/tenants/${OC_TENANT}/graph/social.users/triangles?${qs}`, { headers: { Authorization: `Bearer ${OC_TOKEN}` } }, ); const triangles = await res.json(); ``` #### Go ``` // The Go SDK does not wrap the graph algorithms - use net/http. url := "https://" + OC_HOST + "/v1/tenants/" + OC_TENANT + "/graph/social.users/triangles?rel=follows" req, _ := http.NewRequestWithContext(ctx, "GET", url, nil) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var triangles []map[string][]string json.NewDecoder(resp.Body).Decode(&triangles) ``` ## what you get back ``` [ { "a": ["alice"], "b": ["bob"], "c": ["carol"] }, { "a": ["heidi"], "b": ["ivan"], "c": ["judy"] } ] ``` One object per triangle. `a`, `b` and `c` are primary-key arrays, not strings - the array is what a composite primary key needs, so a single-column key arrives as a one-element array. The three keys come back in canonical order, so the same triangle is always reported the same way and you can deduplicate on the triple. ## how it works - The relation is symmetrised before enumeration, so a triangle is found whether its three edges point around the ring or at each other. - Enumeration runs inside the query executor rather than as a standalone library call, so it inherits the same cancellation and result-budget handling as the rest of the read path. - A self-loop is not an edge for this purpose and contributes no triangle. - On a graph big enough to blow the scan budget the request is refused with a tagged `graph_work_budget_exceeded` 400 instead of being allowed to exhaust memory. ## common mistakes - Expecting direction to matter. The relation is walked as undirected. `alice` pointing at `bob` and `bob` pointing at `alice` are the same edge here. If direction is the question, triangles is the wrong endpoint. - Reading `a` as a string. Each of `a`, `b` and `c` is an array. For a single-column primary key you want `row.a[0]`. - Pointing `rel` at another table. The relation has to target the table it is declared on. A cross-table relation is refused rather than quietly returning nothing. --- # MySQL wire protocol examples - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/mysql Sitemap last modified: 2026-09-13T20:18:50.000Z examples · mysql # MySQL wire protocol examples Four recipes for a client speaking the MySQL wire protocol to an OriginChainDB instance. They were written from a real client driven against the shipping engine - Sequelize 6.37 over mysql2 3.24 - so each one either worked or is marked as refused. The listener is off by default and is enabled per instance. Read [the compatibility reference](https://originchaindb.com/docs/compatibility/mysql-wire) first - it covers TLS, the two credential models, and the limits these recipes stay inside. [1 / 4 Connect and run CRUD Open a TLS session, declare a table in code, then insert, select with a filter and an ordering, update, and delete.](https://originchaindb.com/docs/examples/mysql/connect-and-crud)[2 / 4 Transactions that commit and roll back Both outcomes, exercised against the engine - including the check that proves a rolled-back block really left nothing behind.](https://originchaindb.com/docs/examples/mysql/transactions)[3 / 4 Upsert with ON DUPLICATE KEY UPDATE The MySQL spelling is translated to the engine's upsert. Why REPLACE INTO is refused instead of mapped to it.](https://originchaindb.com/docs/examples/mysql/upsert-on-duplicate-key)[4 / 4 Reflection: what the catalog will answer The information_schema views that work, the SHOW forms that are refused, and how to keep an ORM off the refused half.](https://originchaindb.com/docs/examples/mysql/reflection-and-catalog) MySQL is a registered trademark of Oracle Corporation and/or its affiliates. OriginChainDB is not affiliated with, endorsed by, or sponsored by Oracle Corporation. OriginChainDB implements a compatible wire protocol so that existing MySQL clients can connect to it; it does not distribute MySQL software. --- # Connect and run CRUD - OriginChainDB MySQL example Canonical source: https://originchaindb.com/docs/examples/mysql/connect-and-crud Sitemap last modified: 2026-09-13T20:18:50.000Z examples · mysql · 1 / 4 # 1. Connect and run CRUD [← MySQL examples](https://originchaindb.com/docs/examples/mysql) shared-credential listeners only The writes on this page run only against a listener on the shared credential model. A per-user MySQL session is read-only and refuses every write, so if your tenant has any database RBAC grant or any row-level security policy - the condition that forces per-user login - none of this will run against it. Check which model you are on under [credential models](https://originchaindb.com/docs/compatibility/mysql-wire#credentials) before you build on this, and write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client) if you are not. what this does Opens a TLS session against a listener speaking the MySQL wire protocol, then runs the four operations an application actually spends its day on. This is the workload we measured end to end with Sequelize 6.37 over mysql2 3.24, and all of it worked. the client ``` import mysql2 from "mysql2"; import { Sequelize, DataTypes } from "sequelize"; const sequelize = new Sequelize(process.env.OC_DB, process.env.OC_DB_USER, process.env.OC_DB_PASSWORD, { host: process.env.OC_MYSQL_HOST, port: 3306, dialect: "mysql", dialectModule: mysql2, dialectOptions: { ssl: { minVersion: "TLSv1.2", rejectUnauthorized: true } }, logging: false, }); // Declared in code, on purpose: the listener answers no SHOW form, so an ORM // cannot discover this table for itself. const Order = sequelize.define( "Order", { id: { type: DataTypes.STRING, primaryKey: true }, customer: DataTypes.STRING, total_cents: DataTypes.INTEGER, status: DataTypes.STRING, }, { tableName: "orders", timestamps: false, freezeTableName: true }, ); await sequelize.authenticate(); await Order.create({ id: "o_1", customer: "c_7", total_cents: 4200, status: "open" }); const open = await Order.findAll({ where: { status: "open" }, order: [["total_cents", "DESC"]], limit: 10, }); await Order.update({ status: "shipped" }, { where: { id: "o_1" } }); await Order.destroy({ where: { id: "o_1" } }); ``` the same thing as plain SQL If you are driving the session from a CLI or a thin driver rather than an ORM, this is what the four statements look like. ``` INSERT INTO orders (id, customer, total_cents, status) VALUES ('o_1', 'c_7', 4200, 'open'); SELECT id, customer, total_cents FROM orders WHERE status = 'open' ORDER BY total_cents DESC LIMIT 10; UPDATE orders SET status = 'shipped' WHERE id = 'o_1'; DELETE FROM orders WHERE id = 'o_1'; ``` how it works - Each statement is parsed in MySQL's dialect, rewritten construct by construct into the engine's, and re-rendered before it runs. Backtick-quoted identifiers and all three `LIMIT` spellings are handled on the way through. - The ORM sends these as prepared statements. Placeholders are bound server-side, so values never reach the parser as text. - Row-level security and column masking are applied below the wire, so this session sees exactly what your policies allow for the identity it authenticated as. common mistakes - Calling sync(). `sequelize.sync()` introspects with `SHOW`, which is refused. Create the table with [a schema](https://originchaindb.com/docs/schemas) or a `CREATE TABLE`, and declare it in code. - Leaving TLS to the default. A shared-credential listener with no certificate serves in cleartext without complaining. Ask for TLS in the client and require it on the listener - see [TLS](https://originchaindb.com/docs/compatibility/mysql-wire#tls). - An unbounded DELETE. MySQL's `DELETE … LIMIT` is refused, so bound the statement with `WHERE` instead of a row cap. MySQL is a registered trademark of Oracle Corporation and/or its affiliates. OriginChainDB is not affiliated with, endorsed by, or sponsored by Oracle Corporation. OriginChainDB implements a compatible wire protocol so that existing MySQL clients can connect to it; it does not distribute MySQL software. --- # Reflection and the catalog - OriginChainDB MySQL example Canonical source: https://originchaindb.com/docs/examples/mysql/reflection-and-catalog Sitemap last modified: 2026-09-13T20:18:50.000Z examples · mysql · 4 / 4 # 4. Reflection: what the catalog will answer [← MySQL examples](https://originchaindb.com/docs/examples/mysql) what this does Shows exactly where the metadata boundary sits, because this is the thing most likely to decide whether your tool works. Metadata is split: the information_schema views are answered, and every SHOW form is refused. A tool that reflects through the first half gets somewhere; one that reflects through the second does not. answered Five views, served by the same shared catalog the other wire protocols read: `tables`, `columns`, `table_constraints`, `key_column_usage`, `referential_constraints`. ``` -- List the tables in a schema. SELECT table_name FROM information_schema.tables WHERE table_schema = 'shop'; -- Describe one table's columns. SELECT column_name, data_type, is_nullable FROM information_schema.columns WHERE table_schema = 'shop' AND table_name = 'orders'; -- Constraints, and the columns they cover. SELECT constraint_name, constraint_type FROM information_schema.table_constraints WHERE table_name = 'orders'; SELECT constraint_name, column_name FROM information_schema.key_column_usage WHERE table_name = 'orders'; ``` refused ``` SHOW TABLES; -- refused SHOW DATABASES; -- refused SHOW COLUMNS FROM orders; -- refused SHOW CREATE TABLE orders; -- refused SHOW INDEX FROM orders; -- refused SHOW WARNINGS; -- refused SHOW VARIABLES LIKE 'sql_mode'; -- refused ``` The refusal is deliberate, and it is the safer of the two options. The shared front door carries a compatibility shim for a different protocol that answers any `SHOW x` with a single fabricated row reading `x = 'on'`. Left alone, `SHOW TABLES` would return that row: a confident, wrong answer. Refusing at the dialect boundary turns it into an error you can see. keeping an ORM on the answered half ``` // Safe: you tell the ORM the shape, it never asks the database. const Order = sequelize.define( "Order", { id: { type: DataTypes.STRING, primaryKey: true }, customer: DataTypes.STRING, total_cents: DataTypes.INTEGER, status: DataTypes.STRING, }, { tableName: "orders", timestamps: false, freezeTableName: true }, ); // Not safe on this listener - each of these introspects with SHOW: // await sequelize.sync(); // await sequelize.getQueryInterface().describeTable("orders"); // await sequelize.getQueryInterface().showIndex("orders"); // Also not available: associations. Model the join in your query instead of // asking the ORM to follow a relation for you. // Order.belongsTo(Customer); // will not load ``` decide before you adopt Because reflection is split and associations do not load, this listener suits an application whose tables are declared in code and whose queries touch one table at a time. It does not suit one that models a domain with relations, and no amount of configuration changes that today. If you need relations, use the [SQL endpoint](https://originchaindb.com/docs/sql) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client), both of which carry the full surface today. MySQL is a registered trademark of Oracle Corporation and/or its affiliates. OriginChainDB is not affiliated with, endorsed by, or sponsored by Oracle Corporation. OriginChainDB implements a compatible wire protocol so that existing MySQL clients can connect to it; it does not distribute MySQL software. --- # Transactions - OriginChainDB MySQL example Canonical source: https://originchaindb.com/docs/examples/mysql/transactions Sitemap last modified: 2026-09-13T20:18:50.000Z examples · mysql · 2 / 4 # 2. Transactions that commit and roll back [← MySQL examples](https://originchaindb.com/docs/examples/mysql) shared-credential listeners only The writes on this page run only against a listener on the shared credential model. A per-user MySQL session is read-only and refuses every write, so if your tenant has any database RBAC grant or any row-level security policy - the condition that forces per-user login - none of this will run against it. Check which model you are on under [credential models](https://originchaindb.com/docs/compatibility/mysql-wire#credentials) before you build on this, and write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client) if you are not. what this does Runs both halves of a transaction contract over the wire: a block that commits and one that rolls back. Both were exercised against the shipping engine. The second is the one worth testing yourself, because a transaction layer that only pretends to work still looks correct until something has to be undone. committing ``` const tx = await sequelize.transaction(); try { await Order.create( { id: "o_2", customer: "c_7", total_cents: 900, status: "open" }, { transaction: tx }, ); await Order.update({ status: "paid" }, { where: { id: "o_2" }, transaction: tx }); await tx.commit(); } catch (err) { await tx.rollback(); throw err; } // After the commit the row is there, with status "paid". const saved = await Order.findByPk("o_2"); ``` rolling back ``` const tx = await sequelize.transaction(); await Order.create( { id: "o_3", customer: "c_9", total_cents: 100, status: "open" }, { transaction: tx }, ); await tx.rollback(); // The important half of the test: the row is NOT there afterwards. const gone = await Order.findByPk("o_3"); // null ``` the same thing as plain SQL ``` START TRANSACTION; INSERT INTO orders (id, customer, total_cents, status) VALUES ('o_2', 'c_7', 900, 'open'); UPDATE orders SET status = 'paid' WHERE id = 'o_2'; COMMIT; START TRANSACTION; INSERT INTO orders (id, customer, total_cents, status) VALUES ('o_3', 'c_9', 100, 'open'); ROLLBACK; ``` how it works - A session starts in autocommit, which the handshake reports honestly: with nothing else said, every statement commits on its own. - `START TRANSACTION` opens a real engine transaction, and `COMMIT` / `ROLLBACK` end it for real. - Turning autocommit off is also a real session mode rather than an acknowledged no-op: the engine opens a transaction implicitly around your statements, and switching autocommit back on commits the open block, which is what a MySQL client expects. common mistakes - Forgetting to pass the transaction. A statement issued without `{ transaction: tx }` runs outside the block and will not be rolled back with it. - Batching statements to save a round trip. Multi-statement queries are refused. Send one statement at a time. - Assuming rollback was tested for you. Assert the row is gone, as above. That assertion is the whole value of the test. MySQL is a registered trademark of Oracle Corporation and/or its affiliates. OriginChainDB is not affiliated with, endorsed by, or sponsored by Oracle Corporation. OriginChainDB implements a compatible wire protocol so that existing MySQL clients can connect to it; it does not distribute MySQL software. --- # Upsert with ON DUPLICATE KEY UPDATE - OriginChainDB MySQL example Canonical source: https://originchaindb.com/docs/examples/mysql/upsert-on-duplicate-key Sitemap last modified: 2026-09-13T20:18:50.000Z examples · mysql · 3 / 4 # 3. Upsert with ON DUPLICATE KEY UPDATE [← MySQL examples](https://originchaindb.com/docs/examples/mysql) shared-credential listeners only The writes on this page run only against a listener on the shared credential model. A per-user MySQL session is read-only and refuses every write, so if your tenant has any database RBAC grant or any row-level security policy - the condition that forces per-user login - none of this will run against it. Check which model you are on under [credential models](https://originchaindb.com/docs/compatibility/mysql-wire#credentials) before you build on this, and write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client) if you are not. what this does Inserts a row, or updates the named columns if a row with that key already exists. Write it in the MySQL spelling you already know; the translator rewrites it into the engine's upsert. what you write ``` INSERT INTO orders (id, customer, total_cents, status) VALUES ('o_1', 'c_7', 4200, 'open') ON DUPLICATE KEY UPDATE total_cents = VALUES(total_cents), status = VALUES(status); ``` what the engine runs ``` INSERT INTO orders (id, customer, total_cents, status) VALUES ('o_1', 'c_7', 4200, 'open') ON CONFLICT (id) DO UPDATE SET total_cents = EXCLUDED.total_cents, status = EXCLUDED.status; ``` `VALUES(col)` becomes `EXCLUDED.col` - the value the statement tried to insert. how it works - MySQL's syntax leaves the conflict target implicit - "duplicate key" does not say which key. The engine's upsert needs it named, so the translator resolves it from the catalog and writes it in. - When that resolution is ambiguous it stops instead of guessing. A table carrying a secondary unique index is refused, because picking the wrong arbiter would silently update on the wrong key - a wrong answer rather than an error. On such a table, write the upsert as an explicit `INSERT … ON CONFLICT` through the [SQL endpoint](https://originchaindb.com/docs/sql), where you name the target yourself. why REPLACE INTO is refused ``` mysql> REPLACE INTO orders (id, status) VALUES ('o_1', 'open'); ERROR 1235 (42000): REPLACE INTO is not supported — its delete-then-insert semantics reset unmentioned columns to their defaults, which is not what the engine's upsert does. Use INSERT … ON DUPLICATE KEY UPDATE. ``` `REPLACE INTO` deletes the old row and inserts a new one, so any column you did not mention comes back as its default. The engine's upsert keeps existing values. Mapping one to the other would look like it worked and quietly discard data, so it is refused instead. `INSERT IGNORE` is refused for the same reason: swallowing errors is not a behaviour we will emulate. MySQL is a registered trademark of Oracle Corporation and/or its affiliates. OriginChainDB is not affiliated with, endorsed by, or sponsored by Oracle Corporation. OriginChainDB implements a compatible wire protocol so that existing MySQL clients can connect to it; it does not distribute MySQL software. --- # SQL query examples - SELECT, JOIN, GROUP BY - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/sql Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql # SQL examples [← All examples](https://originchaindb.com/docs/examples) SQL runs over one endpoint - POST /v1/tenants/:t/sql - and comes back as JSON rows keyed by column name. The 13 examples below each carry the schema TOML they need, a pre-seed insert, side-by-side cURL / Python / TypeScript / Go, the response shape, and notes on common mistakes. All 13 execute today. ORDER BY, HAVING, window functions (including explicit ROWS frames), UPDATE, DELETE and both plain and recursive CTEs all run server-side. Each shape has bounds — RANGE frames take peer bounds only, a recursive CTE is UNION ALL and one CTE per statement, and windows cannot yet share a SELECT with a JOIN or GROUP BY. The per-example pages state the limits that apply. [1 SELECT with projection works today Project specific columns from a single table. Read only the fields you need. →](https://originchaindb.com/docs/examples/sql/select-projection)[2 WHERE = on an indexed column works today Equality predicate. The planner promotes it to an IndexScan when an index covers the column. →](https://originchaindb.com/docs/examples/sql/where-equality)[3 WHERE col IN (...) works today Match against a small set of literal values. Folded into a disjunction over equalities. →](https://originchaindb.com/docs/examples/sql/where-in)[4 WHERE col BETWEEN low AND high works today Range predicate on a numeric or string column. Inclusive on both bounds. →](https://originchaindb.com/docs/examples/sql/where-between)[5 INNER JOIN - two tables works today Combine rows that have a match on both sides. Hash-join on the equality. →](https://originchaindb.com/docs/examples/sql/inner-join)[6 LEFT JOIN - keep unmatched left rows works today Every row from the left side, plus matches from the right (null if no match). →](https://originchaindb.com/docs/examples/sql/left-join)[7 GROUP BY + COUNT / SUM / AVG / MIN / MAX works today Multiple aggregates in one pass. The hash aggregator buckets on the GROUP BY key. →](https://originchaindb.com/docs/examples/sql/group-by-aggregate)[8 ORDER BY + LIMIT / OFFSET works today ORDER BY (asc/desc), LIMIT and OFFSET all execute through the SQL translator today. →](https://originchaindb.com/docs/examples/sql/limit-roadmap-orderby)[9 Window functions works today ROW_NUMBER, RANK, LAG and LEAD execute today, and so do explicit ROWS frames on the aggregates. RANGE takes peer bounds only. →](https://originchaindb.com/docs/examples/sql/window-functions-roadmap)[10 Uncorrelated IN (SELECT ...) works today Uncorrelated and correlated subqueries both execute - IN (SELECT), EXISTS / NOT EXISTS, and scalar forms. →](https://originchaindb.com/docs/examples/sql/uncorrelated-in-select)[11 UPDATE via /sql works today UPDATE executes against the engine and returns rows_affected. (Earlier builds only translated.) →](https://originchaindb.com/docs/examples/sql/update-via-sql)[12 DELETE via /sql works today DELETE executes and returns rows_affected. Any supported predicate works, plus RETURNING. →](https://originchaindb.com/docs/examples/sql/delete-via-sql)[13 WITH RECURSIVE works today Recursive CTEs walk a hierarchy to a fixed point: one CTE per statement, UNION ALL, depth and row capped. →](https://originchaindb.com/docs/examples/sql/with-recursive-roadmap) [14 Newer SQL recipes new CTEs, bind parameters, date_trunc time-bucketing, CASE and catalog introspection — the most recent additions, each shown against live output. →](https://originchaindb.com/docs/examples/sql/newer-recipes) --- # DELETE via /sql - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/delete-via-sql Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 12 / 13 · works today # 12. DELETE via /sql [← SQL examples](https://originchaindb.com/docs/examples/sql) works today `DELETE` over `/sql` executes. On its own it removes the matched rows in a single durable frame and answers with `rows_affected`. A `WHERE = ` takes a typed single-row fast path; any other supported predicate - and a composite primary key - lowers to a scan-based delete. Inside `BEGIN; ... COMMIT;` the same statement buffers instead, reporting `rows_buffered` until commit. supported - delete inside a transaction Inside an open transaction, a primary-key `DELETE` buffers a row-removal op and the `COMMIT` applies it durably in one frame. Carry the same session id across all three calls so they share one transaction buffer. #### cURL ``` # Supported SQL route: wrap the DELETE in an explicit transaction. # The same session id must accompany BEGIN, the DELETE, and COMMIT so # the buffered row-removal is applied atomically on COMMIT. SID="sess-$(date +%s)" # 1. open the transaction curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{ "sql": "BEGIN" }' # 2. buffer the delete (matched by primary key) curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{ "sql": "DELETE FROM shop.orders WHERE id = '\''o_003'\''" }' # 3. commit - the buffered row-removal lands durably here curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{ "sql": "COMMIT" }' ``` responses ``` // the DELETE step, inside the open transaction: { "kind": "delete", "schema": "shop.orders", "pk": "o_003", "rows_buffered": 1 } // the COMMIT step - the buffered op is now durable: { "kind": "tx", "op": "commit", "ops_committed": 1 } ``` `rows_buffered` is `1` when the row existed (and `0` when it didn't). The removal is not visible to other readers until `COMMIT` returns `ops_committed: 1`. also supported - typed row delete When you already hold the primary key and don't need SQL, delete the row through the [typed row endpoints](https://originchaindb.com/docs/api#rows). The row delete path carries the same idempotency-key plumbing as the rest of the typed `/rows` surface. also supported - DELETE without a transaction With no `BEGIN` around it, the same statement autocommits: the rows are removed in one frame before the response returns, and `rows_affected` counts them. Foreign keys are enforced on this path too. Note that `/sql` is the one surface that accepts a bare `DELETE FROM t` with no `WHERE` and empties the table - the typed row endpoints refuse that shape. #### cURL ``` # A bare DELETE outside a transaction does NOT reliably remove the row. # The response echoes the translated delete (the pk it resolved) - treat # it as a translation, not a confirmation that the row is gone. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "DELETE FROM shop.orders WHERE id = '\''o_003'\''" }' ``` response (the row is already gone) ``` { "kind": "delete", "schema": "shop.orders", "rows_affected": 1 } ``` Add `RETURNING order_id, total_cents` (or `RETURNING *`) and the response also carries a `rows` array holding the deleted rows' values, projected to that column list. A no-match delete answers `rows_affected: 0` with an empty `rows` array rather than an error. --- # GROUP BY + aggregates - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/group-by-aggregate Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 7 / 13 # 7. GROUP BY + aggregates [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Bucket rows by `status` and compute multiple aggregates per bucket in a single pass - row count, sum, average, min, and max of `amount_cents`. when to use it - Reporting dashboards: "orders per status", "revenue per customer", "rows per day". - You can mix any number of `COUNT` / `SUM` / `AVG` / `MIN` / `MAX` in one query - they all share the same scan. - GROUP BY can list multiple columns: `GROUP BY status, region`. - `DISTINCT` aggregates execute server-side too - `COUNT(DISTINCT customer_id)`, `SUM(DISTINCT amount_cents)`, `AVG(DISTINCT ...)`. - Filter on the aggregated result with `HAVING` - it runs after grouping, server-side. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT status, COUNT(*) AS n, SUM(amount_cents) AS total, AVG(amount_cents) AS avg_amt, MIN(amount_cents) AS min_amt, MAX(amount_cents) AS max_amt FROM shop.orders GROUP BY status" }' ``` #### Python ``` result = db.sql(""" SELECT status, COUNT(*) AS n, SUM(amount_cents) AS total, AVG(amount_cents) AS avg_amt, MIN(amount_cents) AS min_amt, MAX(amount_cents) AS max_amt FROM shop.orders GROUP BY status """) ``` #### TypeScript ``` const result = await db.sql(` SELECT status, COUNT(*) AS n, SUM(amount_cents) AS total, AVG(amount_cents) AS avg_amt, MIN(amount_cents) AS min_amt, MAX(amount_cents) AS max_amt FROM shop.orders GROUP BY status `); ``` #### Go ``` result, _ := db.SQL(ctx, ` SELECT status, COUNT(*) AS n, SUM(amount_cents) AS total, AVG(amount_cents) AS avg_amt, MIN(amount_cents) AS min_amt, MAX(amount_cents) AS max_amt FROM shop.orders GROUP BY status `) ``` what you get back ``` { "kind": "select", "rows": [ { "status": "paid", "n": 2, "total": 6240, "avg_amt": 3120.0, "min_amt": 1250, "max_amt": 4990 }, { "status": "pending", "n": 1, "total": 12900, "avg_amt": 12900.0, "min_amt": 12900, "max_amt": 12900 } ] } ``` common mistakes - Selecting a non-grouped column. Every column in SELECT must either be in GROUP BY or wrapped in an aggregate. Standard SQL rule. - Filtering on aggregates. Use `HAVING` - it executes server-side (e.g. `GROUP BY status HAVING COUNT(*) > 3`). It runs after grouping, so it can reference aggregates that `WHERE` can't. --- # INNER JOIN - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/inner-join Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 5 / 13 # 5. INNER JOIN - two tables [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Combine rows from two tables that share a key. Here, orders + customers joined on `o.customer_id = c.id`. Returns only orders that do have a matching customer. when to use it - You want each row from one table augmented with related data from another. - Both sides should have a match. If you want unmatched left rows too, use [LEFT JOIN](https://originchaindb.com/docs/examples/sql/left-join). - OriginChainDB supports up to 32 joined tables in a single query. the request Uses the `shop.orders` and `shop.customers` schemas from Examples 1 and 3. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT o.id, o.amount_cents, c.email FROM shop.orders o INNER JOIN shop.customers c ON o.customer_id = c.id WHERE o.status = '\''paid'\'' LIMIT 50" }' ``` #### Python ``` result = db.sql(""" SELECT o.id, o.amount_cents, c.email FROM shop.orders o INNER JOIN shop.customers c ON o.customer_id = c.id WHERE o.status = 'paid' LIMIT 50 """) ``` #### TypeScript ``` const result = await db.sql(` SELECT o.id, o.amount_cents, c.email FROM shop.orders o INNER JOIN shop.customers c ON o.customer_id = c.id WHERE o.status = 'paid' LIMIT 50 `); ``` #### Go ``` result, _ := db.SQL(ctx, ` SELECT o.id, o.amount_cents, c.email FROM shop.orders o INNER JOIN shop.customers c ON o.customer_id = c.id WHERE o.status = 'paid' LIMIT 50 `) ``` what you get back ``` { "kind": "select", "rows": [ { "o.id": "o_001", "o.amount_cents": 4990, "c.email": "alice@example.com" }, { "o.id": "o_003", "o.amount_cents": 1250, "c.email": "alice@example.com" } ] } ``` Note that result keys are dotted - `"o.id"`, `"c.email"` - because the engine carries the table alias through to disambiguate same-named columns on both sides. how it works - The planner picks the smaller side (here `customers`) to build a hash table. - It then scans the larger side and probes the hash table for each row. - The WHERE predicate is pushed into the orders-side scan, so only paid orders enter the join. common mistakes - No ON condition. CROSS JOIN isn't supported - every join needs an explicit equi-condition. - Ambiguous column names. If both sides have a column called `id`, qualify with the alias - `o.id`, `c.id`. --- # LEFT JOIN - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/left-join Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 6 / 13 # 6. LEFT JOIN - keep unmatched left rows [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Every row from the left side (customers) comes back, even if no matching row exists on the right (orders). When there's no match, the right-side columns are `null`. In the response below, customer `c_3` has no orders, so `o.id` is null for them. when to use it - You want to keep the parent row even if there are no child rows. Think "all customers, with their orders if any". - To find rows on the left that have no match on the right - add `WHERE o.id IS NULL` after the join. (Note: filtering on the right side of a LEFT JOIN can defeat the LEFT semantics if done wrong - put right-side predicates in the ON clause.) Mirror variants: RIGHT JOIN keeps every row from the right; FULL OUTER JOIN keeps every row from both sides. Both are supported. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT c.id, c.email, o.id FROM shop.customers c LEFT JOIN shop.orders o ON c.id = o.customer_id LIMIT 20" }' ``` #### Python ``` result = db.sql(""" SELECT c.id, c.email, o.id FROM shop.customers c LEFT JOIN shop.orders o ON c.id = o.customer_id LIMIT 20 """) ``` #### TypeScript ``` const result = await db.sql(` SELECT c.id, c.email, o.id FROM shop.customers c LEFT JOIN shop.orders o ON c.id = o.customer_id LIMIT 20 `); ``` #### Go ``` result, _ := db.SQL(ctx, ` SELECT c.id, c.email, o.id FROM shop.customers c LEFT JOIN shop.orders o ON c.id = o.customer_id LIMIT 20 `) ``` what you get back ``` { "kind": "select", "rows": [ { "c.id": "c_1", "c.email": "alice@example.com", "o.id": "o_001" }, { "c.id": "c_1", "c.email": "alice@example.com", "o.id": "o_003" }, { "c.id": "c_2", "c.email": "bob@example.com", "o.id": "o_002" }, { "c.id": "c_3", "c.email": "carol@example.com", "o.id": null } ] } ``` common mistakes - Right-side predicates in WHERE. `WHERE o.status = 'paid'` after a LEFT JOIN turns it back into an INNER JOIN (because the unmatched rows have null status, which fails the predicate). Put right-side predicates in the ON clause if you want LEFT semantics preserved. - Confusing LEFT and RIGHT. If you find yourself reaching for RIGHT JOIN, swap the table order and use LEFT - it reads more clearly. --- # ORDER BY + LIMIT / OFFSET - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/limit-roadmap-orderby Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 8 / 13 · works today # 8. ORDER BY + LIMIT / OFFSET [← SQL examples](https://originchaindb.com/docs/examples/sql) works today `ORDER BY` (ascending or descending), `LIMIT`, and `OFFSET` all execute server-side through the SQL translator - the engine's `Sort` operator is wired through. Sorting and pagination happen in the engine, not your app. top 3 by amount Uses the same `shop.orders` schema and seed from [Example 3](https://originchaindb.com/docs/examples/sql/where-in). #### cURL ``` # ORDER BY + LIMIT + OFFSET all execute server-side. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, amount_cents FROM shop.orders WHERE amount_cents > 0 ORDER BY amount_cents DESC LIMIT 3" }' ``` pagination - LIMIT + OFFSET Stable pagination over a sorted result: page N is `ORDER BY ... LIMIT n OFFSET (N-1)*n`. #### cURL ``` # Page 2 (rows 11-20): ORDER BY + LIMIT + OFFSET. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, amount_cents FROM shop.orders ORDER BY amount_cents DESC LIMIT 10 OFFSET 10" }' ``` --- # Newer SQL recipes - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/newer-recipes Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · newer capabilities # Newer SQL recipes [← SQL examples](https://originchaindb.com/docs/examples/sql) verified on the live engine CTEs (recursive and not), positional bind parameters, `date_trunc` time-bucketing, `CASE` expressions and catalog introspection all execute over `/sql`. Every request and every result row on this page was captured from a running engine — not aspirational. For the full supported surface see [SQL over HTTP](https://originchaindb.com/docs/sql); for running totals with `ROWS` / `RANGE` frames see [Window functions](https://originchaindb.com/docs/examples/sql/window-functions-roadmap). CTEs · WITH … AS ## Name a subquery, then aggregate over it A non-recursive `WITH` block names an intermediate result you can filter, join and aggregate — all resolved server-side in one request. The CTE body reads a real table. request · SQL ``` WITH paid AS ( SELECT customer_id, amount_cents FROM shop.orders WHERE status = 'paid' ) SELECT customer_id, count(*) AS n, sum(amount_cents) AS total FROM paid GROUP BY customer_id ``` result ``` { "kind": "select", "columns": ["customer_id", "n", "total"], "rows": [ { "customer_id": 1, "n": 2, "total": 18800 }, { "customer_id": 2, "n": 1, "total": 22000 } ] } ``` Recursive CTEs (`WITH RECURSIVE`) execute over `/sql` too, under their own bounds — one recursive CTE per statement, `UNION ALL` between the arms, and depth and row caps on the walk. See [WITH RECURSIVE](https://originchaindb.com/docs/examples/sql/with-recursive-roadmap). A shape outside those bounds comes back as a clear `400` naming the constraint it broke, never a silent re-interpretation. Parameters · $1 ## Bind values with a params array Keep SQL and values apart: write positional `$1`, `$2` placeholders and pass a `params` array alongside the statement. No string interpolation, no injection surface. request · JSON body ``` { "sql": "SELECT id, amount_cents, status FROM shop.orders WHERE customer_id = $1 AND amount_cents > $2 ORDER BY amount_cents DESC", "params": [1, 1000] } ``` result ``` { "kind": "select", "columns": ["id", "amount_cents", "status"], "rows": [ { "id": 1, "amount_cents": 14900, "status": "paid" }, { "id": 4, "amount_cents": 3900, "status": "paid" } ] } ``` Bind parameters are also how you write timestamps today — pass the value in `params` rather than as an inline `TIMESTAMP '…'` literal: { "sql": "INSERT INTO shop.orders (customer_id, amount_cents, status, created_at) VALUES ($1,$2,$3,$4)", "params": [1, 14900, "paid", "2026-01-04 10:00:00"] } → `{"inserted": 1}`. Time-bucketing · date_trunc ## Roll rows up into time buckets The dashboard shape: bucket a timestamp column by `day` / `month` and aggregate per bucket, in a single `GROUP BY`. request · SQL ``` SELECT date_trunc('month', created_at) AS m, count(*) AS n, sum(amount_cents) AS rev FROM shop.orders GROUP BY date_trunc('month', created_at) ``` result ``` { "kind": "select", "columns": ["m", "n", "rev"], "rows": [ { "m": 1767225600000, "n": 2, "rev": 18800 }, { "m": 1769904000000, "n": 2, "rev": 29200 } ] } ``` Buckets come back as epoch-milliseconds (`1767225600000` = 2026-01-01). Sort the buckets in your app for now — `ORDER BY` over a `date_trunc` group key isn't wired through `/sql` yet. CASE ## Branch inline with CASE Derive a column from a per-row condition without a second query — `CASE WHEN … THEN … ELSE … END` evaluates in the SELECT list. request · SQL ``` SELECT id, amount_cents, CASE WHEN amount_cents >= 10000 THEN 'big' ELSE 'small' END AS tier FROM shop.orders ORDER BY id ``` result ``` { "kind": "select", "columns": ["id", "amount_cents", "tier"], "rows": [ { "id": 1, "amount_cents": 14900, "tier": "big" }, { "id": 3, "amount_cents": 7200, "tier": "small" }, { "id": 4, "amount_cents": 3900, "tier": "small" }, { "id": 6, "amount_cents": 22000, "tier": "big" } ] } ``` Catalog & session ## Ask the engine about itself Standard Postgres introspection works — read `information_schema` / `pg_catalog`, and inspect session state with `SHOW` and `current_schema()`. This is what lets ORMs and SQL clients introspect a schema. request · SQL ``` SELECT column_name FROM information_schema.columns WHERE table_name = 'orders' ``` result ``` { "kind": "select", "columns": ["column_name"], "rows": [ { "column_name": "id" }, { "column_name": "customer_id" }, { "column_name": "amount_cents" }, { "column_name": "status" }, { "column_name": "created_at" } ] } ``` `SHOW search_path` → `public` · `SELECT current_schema()` → `public`. `SET search_path` and GUCs persist for the session. --- # SELECT with projection - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/select-projection Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 1 / 13 # 1. SELECT with projection [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Read specific columns from one table. We ask for just `id` and `email` from the customers table - the engine reads only those columns from the row payload, not the whole row. when to use it - You only need a couple of fields per row, not the whole record. - You're driving a UI list view (id + display field) and the rest of the row would just be bytes on the wire. Use `SELECT *` when you genuinely need every column - explicit projection is just an optimization, not a correctness requirement. the schema Register this once with `POST /v1/tenants/:t/schemas` (Content-Type: `text/plain`). ``` namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "email" ty = "str" [[columns]] name = "region" ty = "str" [[columns]] name = "tier" ty = "str" [[indexes]] name = "by_region" columns = ["region"] ``` seed data Load three customers via the bulk-insert endpoint so the response below has something to return. ``` # One-time seed of three customers - skip if you've already loaded data. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.customers/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "c_1", "email": "alice@example.com", "region": "IN", "tier": "gold" }, { "id": "c_2", "email": "bob@example.com", "region": "US", "tier": "silver" }, { "id": "c_3", "email": "carol@example.com", "region": "DE", "tier": "gold" } ]' ``` the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, email FROM shop.customers LIMIT 5" }' ``` #### Python ``` result = db.sql("SELECT id, email FROM shop.customers LIMIT 5") for row in result.rows: print(row["id"], row["email"]) ``` #### TypeScript ``` const result = await db.sql("SELECT id, email FROM shop.customers LIMIT 5"); if (result.kind === "select") { for (const row of result.rows) { console.log(row.id, row.email); } } ``` #### Go ``` result, err := db.SQL(ctx, "SELECT id, email FROM shop.customers LIMIT 5") if err != nil { /* handle */ } for _, row := range result.Rows { fmt.Println(row["id"], row["email"]) } ``` what you get back ``` { "kind": "select", "rows": [ { "id": "c_1", "email": "alice@example.com" }, { "id": "c_2", "email": "bob@example.com" }, { "id": "c_3", "email": "carol@example.com" } ] } ``` Every `/sql` response carries a `kind` field. `"select"` means the engine ran the query and the rows are in `rows`; each row is a JSON object keyed by column name. how it works - The SQL parser turns the statement into an AST. - The translator folds the AST into a Plan tree - here, `Limit(5) → Project(id, email) → Scan(shop.customers)`. - The executor runs the plan bottom-up. `Scan` iterates the row store; `Project` drops every field except the ones you asked for; `Limit` stops at 5 rows. - The trimmed rows come back as JSON. common mistakes - Missing LIMIT. Without a LIMIT, the query returns every row. On large tables that's slow and expensive. Add LIMIT early, drop it only when you mean it. - Forgot to qualify the table. Tables are addressed as `namespace.table` - `FROM shop.customers`, not `FROM customers`. - Asking for an undeclared column. Selecting a column that doesn't exist in the schema returns `400 unknown_column`. --- # Uncorrelated IN (SELECT ...) - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/uncorrelated-in-select Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 10 / 13 # 10. Uncorrelated IN (SELECT ...) [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Two-step query: the inner SELECT returns the list of customer IDs in region `'IN'`; the outer SELECT returns orders whose `customer_id` is in that list. This is the uncorrelated shape - the inner query doesn't reference any column from the outer query, so it runs once and its result is reused. when to use it - "Find rows in table A where some column matches a result set from table B". - The inner result is small enough to fit in memory (the engine materializes it). - You could also write this as an INNER JOIN with SELECT DISTINCT - they're equivalent. Pick whichever reads more naturally. the request #### cURL ``` # Orders for customers in the 'IN' region. # The inner SELECT is uncorrelated - it doesn't reference outer columns. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, amount_cents FROM shop.orders WHERE customer_id IN (SELECT id FROM shop.customers WHERE region = '\''IN'\'')" }' ``` #### Python ``` result = db.sql(""" SELECT id, amount_cents FROM shop.orders WHERE customer_id IN ( SELECT id FROM shop.customers WHERE region = 'IN' ) """) ``` #### TypeScript ``` const result = await db.sql(` SELECT id, amount_cents FROM shop.orders WHERE customer_id IN ( SELECT id FROM shop.customers WHERE region = 'IN' ) `); ``` #### Go ``` result, _ := db.SQL(ctx, ` SELECT id, amount_cents FROM shop.orders WHERE customer_id IN ( SELECT id FROM shop.customers WHERE region = 'IN' ) `) ``` what you get back ``` { "kind": "select", "rows": [ { "id": "o_001", "amount_cents": 4990 }, { "id": "o_003", "amount_cents": 1250 } ] } ``` correlated subqueries also execute This page shows the uncorrelated shape (the inner runs once, result inlined). Correlated subqueries - where the inner SELECT references an outer column - also execute server-side, evaluated per outer row. You don't have to flatten them to a `JOIN` (though a JOIN is often the faster plan for large inputs). - Correlated EXISTS / NOT EXISTS. `WHERE EXISTS (SELECT 1 FROM x WHERE x.id = o.x_id)` - runs as a semi-join (anti-join for NOT EXISTS). - Correlated IN. `WHERE col IN (SELECT k FROM y WHERE y.j = o.j)` - evaluated per outer row. - Scalar correlated subqueries. `WHERE col = (SELECT COUNT(*) FROM y WHERE y.k = o.k)` - compared against the per-row scalar. --- # UPDATE via /sql - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/update-via-sql Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 11 / 13 · works today # 11. UPDATE via /sql [← SQL examples](https://originchaindb.com/docs/examples/sql) works today `UPDATE` through `/sql` executes against the engine and returns `rows_affected`. (Earlier builds only translated the statement - that's no longer the case.) the /sql path #### cURL ``` # UPDATE through /sql executes against the engine. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "UPDATE shop.orders SET status = '\''shipped'\'' WHERE id = '\''o_001'\''" }' ``` response ``` { "kind": "update", "schema": "shop.orders", "rows_affected": 1 } ``` The matched row is mutated in place. `rows_affected` is 0 when the `WHERE` matches nothing. scope - single row by primary key `UPDATE` via `/sql` targets one row identified by its primary key - the `WHERE` must pin the PK (`WHERE id = '...'`). Set-based updates that match on a non-PK column (e.g. `UPDATE shop.orders SET status = 'shipped' WHERE customer_id = 'c_1'`) are not applied as a bulk mutation. To update many rows, read the matching primary keys first, then issue one PK-pinned `UPDATE` - or one `PUT` to the [row endpoint](https://originchaindb.com/docs/api#rows) - per key. alternative - full-row PUT When you already have the complete new row, a `PUT` to the row endpoint overwrites it atomically in one round-trip and carries the SDK's idempotency-key plumbing. #### cURL ``` # Alternative: PUT the full row representation directly. # Use this when you want to overwrite the whole row in one shot. curl -X PUT "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/o_001" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "o_001", "customer_id": "c_1", "amount_cents": 4990, "status": "shipped" }' ``` --- # WHERE BETWEEN - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/where-between Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 4 / 13 # 4. WHERE col BETWEEN low AND high [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Match rows whose column value falls inside a range. Both bounds are inclusive: `BETWEEN 1000 AND 10000` is the same as `>= 1000 AND <= 10000`. when to use it - Price ranges, age ranges, date ranges (with timestamps stored as `u64` epoch ms). - Anywhere you'd otherwise write `col >= X AND col <= Y`. the request Uses the same `shop.orders` schema and seed from [Example 3](https://originchaindb.com/docs/examples/sql/where-in). #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, amount_cents FROM shop.orders WHERE amount_cents BETWEEN 1000 AND 10000 LIMIT 50" }' ``` #### Python ``` result = db.sql(""" SELECT id, amount_cents FROM shop.orders WHERE amount_cents BETWEEN 1000 AND 10000 LIMIT 50 """) ``` #### TypeScript ``` const result = await db.sql(` SELECT id, amount_cents FROM shop.orders WHERE amount_cents BETWEEN 1000 AND 10000 LIMIT 50 `); ``` #### Go ``` result, _ := db.SQL(ctx, ` SELECT id, amount_cents FROM shop.orders WHERE amount_cents BETWEEN 1000 AND 10000 LIMIT 50 `) ``` what you get back ``` { "kind": "select", "rows": [ { "id": "o_001", "amount_cents": 4990 }, { "id": "o_003", "amount_cents": 1250 } ] } ``` common mistakes - Off-by-one. BETWEEN is inclusive on both ends. If you mean exclusive, use `col > low AND col < high`. - Money as floats. Always store money as integer minor units (`i64` cents). Floats accumulate rounding error - never reliable for range predicates on money. --- # WHERE equality - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/where-equality Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 2 / 13 # 2. WHERE = (equality filter) [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Return only the rows whose `region` column exactly equals the string `'IN'`. Because we declared an index on `region` in the schema, the planner promotes this to an `IndexScan` - the matching row is found by direct lookup rather than by scanning every row. when to use it - "Find the row(s) where this column equals this value" - the most common filter shape. - Lookups by user ID, status, region, category - any single equality predicate. - The column you're filtering on should be either the primary key or covered by a secondary index. Without an index, equality still works but becomes a full scan. the request Same schema and seed data as [Example 1](https://originchaindb.com/docs/examples/sql/select-projection) - if you haven't loaded the customers yet, set those up first. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, email, region FROM shop.customers WHERE region = '\''IN'\'' LIMIT 10" }' ``` #### Python ``` result = db.sql(""" SELECT id, email, region FROM shop.customers WHERE region = 'IN' LIMIT 10 """) for row in result.rows: print(row["id"], row["email"], row["region"]) ``` #### TypeScript ``` const result = await db.sql(` SELECT id, email, region FROM shop.customers WHERE region = 'IN' LIMIT 10 `); if (result.kind === "select") { for (const row of result.rows) { console.log(row.id, row.email, row.region); } } ``` #### Go ``` result, _ := db.SQL(ctx, ` SELECT id, email, region FROM shop.customers WHERE region = 'IN' LIMIT 10 `) for _, row := range result.Rows { fmt.Println(row["id"], row["email"], row["region"]) } ``` what you get back ``` { "kind": "select", "rows": [ { "id": "c_1", "email": "alice@example.com", "region": "IN" } ] } ``` how it works - The translator builds a `Filter(region = 'IN') → Scan(shop.customers)` plan. - The optimiser checks the schema for indexes covering the filtered column. `by_region` covers `region`, so the plan is rewritten to `IndexScan(by_region, region = 'IN')`. - The IndexScan goes directly to the matching keys instead of iterating the whole row store. Lookup is sub-linear instead of linear in row count. Run `EXPLAIN SELECT ...` on the same query to see whether your indexes are being used. common mistakes - Quoting strings as double-quoted. SQL strings use single quotes - `WHERE region = 'IN'`. Double-quoted identifiers refer to columns, not strings. - Case sensitivity. String equality is case-sensitive: `'IN'` doesn't match `'in'`. Lowercase at write time if you want case-insensitive lookups. - Combining with OR. The engine only supports AND in WHERE today. For OR, rewrite as `IN (...)` (next example) or run two queries. --- # WHERE IN value lists - OriginChainDB SQL example Canonical source: https://originchaindb.com/docs/examples/sql/where-in Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 3 / 13 # 3. WHERE col IN (...) [← SQL examples](https://originchaindb.com/docs/examples/sql) what this does Return rows whose `status` matches any value in a small set of literals. `IN (...)` is the OriginChainDB-supported way to express what `x = 'pending' OR x = 'paid'` would mean in traditional SQL - OR in WHERE isn't supported yet, so use IN for value alternatives. when to use it - "Match one of these values" - small enumerations, status fields, category lists. - As a replacement for OR-of-equalities. - The IN list is short (under ~100 values). For larger sets, consider a JOIN against a temp table. the schema ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer_id" ty = "str" [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[indexes]] name = "by_status" columns = ["status"] ``` seed data ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "o_001", "customer_id": "c_1", "amount_cents": 4990, "status": "paid" }, { "id": "o_002", "customer_id": "c_2", "amount_cents": 12900, "status": "pending" }, { "id": "o_003", "customer_id": "c_1", "amount_cents": 1250, "status": "paid" } ]' ``` the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, status FROM shop.orders WHERE status IN ('\''pending'\'', '\''paid'\'') LIMIT 50" }' ``` #### Python ``` result = db.sql(""" SELECT id, status FROM shop.orders WHERE status IN ('pending', 'paid') LIMIT 50 """) for row in result.rows: print(row["id"], row["status"]) ``` #### TypeScript ``` const result = await db.sql(` SELECT id, status FROM shop.orders WHERE status IN ('pending', 'paid') LIMIT 50 `); if (result.kind === "select") { for (const row of result.rows) { console.log(row.id, row.status); } } ``` #### Go ``` result, _ := db.SQL(ctx, ` SELECT id, status FROM shop.orders WHERE status IN ('pending', 'paid') LIMIT 50 `) for _, row := range result.Rows { fmt.Println(row["id"], row["status"]) } ``` what you get back ``` { "kind": "select", "rows": [ { "id": "o_001", "status": "paid" }, { "id": "o_002", "status": "pending" }, { "id": "o_003", "status": "paid" } ] } ``` how it works The translator folds `status IN ('pending', 'paid')` into a disjunction over equalities internally and dispatches the index lookups in parallel. From the engine's perspective each candidate value is treated like a separate IndexScan and the results are merged. common mistakes - Very long IN lists. A few hundred values is fine. Thousands? Bulk-insert the list into a temp table and JOIN. The engine accepts long IN lists but the parser overhead grows with size. - NOT IN with NULL. If any value in the list is NULL, NOT IN returns no rows (standard SQL three-valued logic). Filter NULL explicitly before negating. --- # Window functions - ROW_NUMBER, RANK - OriginChainDB SQL Canonical source: https://originchaindb.com/docs/examples/sql/window-functions-roadmap Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 9 / 13 · works today # 9. Window functions [← SQL examples](https://originchaindb.com/docs/examples/sql) works today `ROW_NUMBER()`, `RANK()`, `DENSE_RANK()`, `LAG()` and `LEAD()` with `OVER (PARTITION BY ... ORDER BY ...)` all execute server-side through the SQL translator, alongside `SUM` / `AVG` / `COUNT` / `MIN` / `MAX` over a window and the value and distribution functions (`FIRST_VALUE`, `LAST_VALUE`, `NTH_VALUE`, `NTILE`, `PERCENT_RANK`, `CUME_DIST`). Explicit `ROWS BETWEEN` / `RANGE` frame clauses (running totals, moving windows) execute too — on the frame-respecting functions: the aggregates plus `FIRST_VALUE` / `LAST_VALUE` / `NTH_VALUE`. Ranking and offset functions ignore a frame by definition, so writing one on `ROW_NUMBER`, `RANK`, `LAG`, `LEAD` or `NTILE` is refused with a hint rather than silently dropped. per-group ranking Rank each customer's orders by amount - the canonical window use - runs server-side: ``` SELECT id, customer_id, amount_cents, ROW_NUMBER() OVER ( PARTITION BY customer_id ORDER BY amount_cents DESC ) AS rn FROM shop.orders ``` Each row comes back with its `rn` computed in one pass - no client-side sorting needed. running totals — ROWS frame Add a `ROWS BETWEEN` frame to accumulate a running total per partition — the classic dashboard shape — in one server-side pass: ``` SELECT id, customer_id, amount_cents, SUM(amount_cents) OVER ( PARTITION BY customer_id ORDER BY id ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW ) AS running FROM shop.orders ``` `running` carries the cumulative sum as the window slides — customer 1 → 14900, 18800; customer 2 → 7200, 14400, 36400. The two frame units differ in what they accept. `ROWS` takes every bound — `UNBOUNDED PRECEDING`, `N PRECEDING`, `CURRENT ROW`, `N FOLLOWING`, `UNBOUNDED FOLLOWING` — so a moving window such as `ROWS BETWEEN 1 PRECEDING AND CURRENT ROW` works. `RANGE` takes peer bounds only (`UNBOUNDED PRECEDING`, `CURRENT ROW`, `UNBOUNDED FOLLOWING`), where `CURRENT ROW` spans the whole tie group; a value offset such as `RANGE BETWEEN 5 PRECEDING` is refused, as are the `GROUPS` frame unit and `EXCLUDE`. what a window cannot share a SELECT with A window function runs over a single-table SELECT. It cannot yet appear in the same `SELECT` as a `JOIN`, a `GROUP BY` or `DISTINCT`, and the `OVER` call has to sit in the select list - you can `ORDER BY` the alias it was given, but not write the call itself into a `WHERE` or `ORDER BY` expression. Do the join or the aggregate in a subquery first, then window over its output. Named windows (`WINDOW w AS (...)` with `OVER w`) are supported; chaining one named window off another is not. --- # WITH RECURSIVE - recursive CTEs - OriginChainDB SQL Canonical source: https://originchaindb.com/docs/examples/sql/with-recursive-roadmap Sitemap last modified: 2026-09-13T20:18:50.000Z examples · sql · 13 / 13 · works today # 13. WITH RECURSIVE [← SQL examples](https://originchaindb.com/docs/examples/sql) works today Recursive CTEs (`WITH RECURSIVE ...`) execute over `/sql`. There is no prefix gate on `WITH` - a statement that starts with it routes through the same path a `SELECT` takes, and plain non-recursive CTEs run there too. The recursive form is evaluated to a fixed point and returns real rows. Bounds worth knowing before you use it are below, and for wide hierarchy walks the graph endpoints are still the better tool. the classic case Walk an org chart down from one root. This runs as written: ``` WITH RECURSIVE descendants AS ( SELECT id, manager_id FROM hr.employees WHERE id = 'ceo' UNION ALL SELECT e.id, e.manager_id FROM hr.employees e INNER JOIN descendants d ON e.manager_id = d.id ) SELECT * FROM descendants ``` The response carries the whole transitive closure from the seed row, and nothing outside it. Access control covers the real tables named in both arms, so a seed or step term you cannot read is refused rather than silently skipped. the bounds - `UNION ALL` between the seed and step arms. `UNION` (distinct) is refused - deduplicating would mean hashing every intermediate row. - Exactly one CTE in a statement that uses `RECURSIVE`; chaining a second binding off the recursive one is refused. - No per-CTE column list (`WITH RECURSIVE t(a, b) AS ...`). - No aggregate over the recursive relation in the outer projection - `SELECT COUNT(*) FROM t` is refused, so count client-side. (Over a plain CTE, `COUNT(*)` does fold - the asymmetry is deliberate.) - Two caps bound the walk: a depth cap of 100 iterations, and a separate cap of 1,000,000 accumulated rows. Cyclic or wide-branching data hits one of them and comes back as a `400` naming the cap it hit - it will not spin or exhaust the engine. Add a visited-set predicate to make such a walk terminate with rows instead. - Filtering and projecting the recursive relation in the outer query works: `SELECT id FROM descendants WHERE id <> 'ceo'`. the alternative - graph BFS For a walk that branches widely, or one you want to run repeatedly, declare the parent column as a relation on the schema and walk the edge with depth-bounded BFS instead. This is not a workaround for a missing feature - it is the traversal-shaped tool for a traversal-shaped job. ``` # A hierarchy walk via the graph endpoint instead of WITH RECURSIVE. # Assumes the schema declares manager_id as a [[relations]] block: # # [[relations]] # name = "reports_to" # from_col = "manager_id" # bidirectional = true # # [relations.target] # namespace = "hr" # table = "employees" # pk = "id" curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/hr.employees/bfs?rel=reports_to&pk=ceo&max_depth=10" \ -H "Authorization: Bearer $OC_TOKEN" ``` See [Graph reference](https://originchaindb.com/docs/graph) for BFS, shortest path, k-shortest, and the other 16 graph algorithms. [Schema reference → relations](https://originchaindb.com/docs/schemas/reference#relations) covers how to declare the relation. when to prefer graph - BFS takes `max_depth` as an explicit per-request argument. `WITH RECURSIVE` is capped too, but by fixed engine limits you tune the query around rather than pass in. - Forward + reverse traversal are both indexed (bidirectional relations) - "who reports to X" and "who does Y report to" are equally fast. - For weighted paths or shortest-path, you can swap BFS for Dijkstra or k-shortest without rewriting the schema. --- # SQL Server wire protocol examples - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/sqlserver Sitemap last modified: 2026-09-13T20:18:50.000Z examples · tds # SQL Server wire examples [← All examples](https://originchaindb.com/docs/examples) Five worked examples against OriginChainDB's TDS listener — the wire protocol used by Microsoft SQL Server. The two connection examples are shown at the driver versions they were actually driven at; the two dialect examples show translation and refusal side by side, because the second one is what decides whether a migration is small or large. Read [the TDS reference](https://originchaindb.com/docs/compatibility/sqlserver-tds) first — in particular the T-SQL boundary. The listener is a preview surface, is off unless someone enabled it on your instance, and requires TLS to exist at all. [1 Connect with mssql-jdbc measured A complete JDBC connection string, the TLS settings that are not optional, and a first query that returns rows. →](https://originchaindb.com/docs/examples/sqlserver/connect-jdbc)[2 Connect with microsoft/go-mssqldb measured The same connection over the Go driver, built from a URL so credentials never land in a hand-concatenated string. →](https://originchaindb.com/docs/examples/sqlserver/connect-go)[3 Prepared statements and batches measured Prepare once, execute many, close the handle - the driver-issued RPC path, and why it is not stored-procedure support. →](https://originchaindb.com/docs/examples/sqlserver/prepared-statements)[4 T-SQL that translates translated Bracket identifiers, TOP, OFFSET/FETCH and the scalar-function set, each shown beside the SQL it becomes. →](https://originchaindb.com/docs/examples/sqlserver/tsql-translated)[5 T-SQL that is refused refused Variables, temp tables, GO, MERGE and date arithmetic - what comes back instead, and the portable rewrite. →](https://originchaindb.com/docs/examples/sqlserver/tsql-refused) Microsoft, SQL Server and T-SQL are trademarks of the Microsoft group of companies; mssql-jdbc and go-mssqldb are Microsoft's own drivers. They are named here only to describe protocol and dialect compatibility and to identify the versions it was measured against. OriginChainDB is not affiliated with, endorsed by, or sponsored by Microsoft, and does not distribute Microsoft software. --- # Connect with go-mssqldb - OriginChainDB TDS example Canonical source: https://originchaindb.com/docs/examples/sqlserver/connect-go Sitemap last modified: 2026-09-13T20:18:50.000Z examples · tds · 2 / 5 # 2. Connect with microsoft/go-mssqldb [← SQL Server wire examples](https://originchaindb.com/docs/examples/sqlserver) what this does The same session as the JDBC example, over Microsoft's Go driver, with the DSN built through `net/url` so a password containing a reserved character cannot quietly corrupt it. Measured against microsoft/go-mssqldb 1.11.0, 21 of 21 checks, on 2026-09-07. the code ``` package main import ( "database/sql" "fmt" "net/url" "os" _ "github.com/microsoft/go-mssqldb" ) func dsn() string { q := url.Values{} q.Set("database", os.Getenv("OC_DATABASE")) q.Set("encrypt", "true") // mandatory - the only way in q.Set("TrustServerCertificate", "false") // validate the chain q.Set("dial timeout", "30") u := url.URL{ Scheme: "sqlserver", // url.UserPassword escapes for you. Never build this by hand. User: url.UserPassword(os.Getenv("OC_DB_USER"), os.Getenv("OC_DB_PASSWORD")), Host: os.Getenv("OC_SQL_HOST") + ":1433", RawQuery: q.Encode(), } return u.String() } func main() { db, err := sql.Open("sqlserver", dsn()) if err != nil { panic(err) } defer db.Close() // sql.Open is lazy. Ping is the first thing that actually connects. if err := db.Ping(); err != nil { panic(err) } rows, err := db.Query( "SELECT TOP 5 id, customer_ref, amount_cents"+ " FROM invoices WHERE status = @p1 ORDER BY amount_cents DESC", "open") if err != nil { panic(err) } defer rows.Close() for rows.Next() { var id, cents int var ref string if err := rows.Scan(&id, &ref, ¢s); err != nil { panic(err) } fmt.Printf("%d %-12s %d\n", id, ref, cents) } if err := rows.Err(); err != nil { panic(err) } } ``` the result shape ``` 4101 ACME-77 129900 4088 ACME-12 98450 4077 ORBIT-3 61000 4051 ACME-77 44900 4020 ORBIT-9 12750 ``` notes - `sql.Open` does not connect. If a listener is missing or a certificate is unloadable you will not learn it until `Ping` or the first query, so call `Ping` at start-up rather than discovering it inside a request. - This driver names parameters `@p1`, `@p2`, and so on. That is the driver's placeholder syntax and it is fine — it is rewritten to a positional bind before the statement reaches the engine. It is not the same thing as a T-SQL `@variable`, which is [refused](https://originchaindb.com/docs/examples/sqlserver/tsql-refused). - Leave `TrustServerCertificate` at `false`. Turning it on does not make a broken listener work — there is no listener to reach when the certificate is the problem — it only stops you noticing who you are talking to. Microsoft, SQL Server and T-SQL are trademarks of the Microsoft group of companies; go-mssqldb is Microsoft's own driver. Named here only to describe protocol compatibility and to identify the version it was measured against. OriginChainDB is not affiliated with, endorsed by, or sponsored by Microsoft, and does not distribute Microsoft software. --- # Connect with mssql-jdbc - OriginChainDB TDS example Canonical source: https://originchaindb.com/docs/examples/sqlserver/connect-jdbc Sitemap last modified: 2026-09-13T20:18:50.000Z examples · tds · 1 / 5 # 1. Connect with mssql-jdbc [← SQL Server wire examples](https://originchaindb.com/docs/examples/sqlserver) what this does Opens a session against an OriginChainDB TDS listener with Microsoft's JDBC driver, and reads five rows back. Measured against mssql-jdbc 13.4.0 and 12.10.1, both 34 of 34 checks, on 2026-09-07. the environment ``` export OC_SQL_HOST=... # the endpoint your console shows export OC_DATABASE=... export OC_DB_USER=... export OC_DB_PASSWORD=... # never inline this in the URL ``` the code ``` import java.sql.*; public class Connect { public static void main(String[] args) throws Exception { String url = "jdbc:sqlserver://" + System.getenv("OC_SQL_HOST") + ":1433" + ";databaseName=" + System.getenv("OC_DATABASE") // TLS is mandatory. Without it there is no listener to reach. + ";encrypt=true" // Validate the chain. Only relax this against a test instance. + ";trustServerCertificate=false" + ";hostNameInCertificate=" + System.getenv("OC_SQL_HOST") + ";loginTimeout=30"; try (Connection conn = DriverManager.getConnection( url, System.getenv("OC_DB_USER"), System.getenv("OC_DB_PASSWORD")); Statement st = conn.createStatement(); ResultSet rs = st.executeQuery( "SELECT TOP 5 id, customer_ref, amount_cents" + " FROM invoices ORDER BY amount_cents DESC")) { while (rs.next()) { System.out.printf("%d %-12s %d%n", rs.getInt("id"), rs.getString("customer_ref"), rs.getInt("amount_cents")); } } } } ``` the result shape ``` 4101 ACME-77 129900 4088 ACME-12 98450 4077 ORBIT-3 61000 4051 ACME-77 44900 4020 ORBIT-9 12750 ``` notes - `encrypt=true` is not a hardening option here, it is the only way in. A listener with no loadable certificate never binds, so a certificate problem shows up as a connection that finds nothing at all: ``` com.microsoft.sqlserver.jdbc.SQLServerException: The TCP/IP connection to the host ..., port 1433 has failed. ``` - If that is what you see, check the certificate before the firewall. Then check whether a listener was ever enabled — it is off by default on every instance. - `TOP 5` is translated to a `LIMIT`. The [translated set](https://originchaindb.com/docs/examples/sqlserver/tsql-translated) lists what else is rewritten, and the [refused set](https://originchaindb.com/docs/examples/sqlserver/tsql-refused) lists what is not. - Keep the password out of the URL. The driver takes it as a separate argument, and a URL is the thing that ends up in a log. Microsoft, SQL Server and T-SQL are trademarks of the Microsoft group of companies; mssql-jdbc is Microsoft's own driver. Named here only to describe protocol compatibility and to identify the versions it was measured against. OriginChainDB is not affiliated with, endorsed by, or sponsored by Microsoft, and does not distribute Microsoft software. --- # Prepared statements over TDS - OriginChainDB example Canonical source: https://originchaindb.com/docs/examples/sqlserver/prepared-statements Sitemap last modified: 2026-09-13T20:18:50.000Z examples · tds · 3 / 5 # 3. Prepared statements and batches [← SQL Server wire examples](https://originchaindb.com/docs/examples/sqlserver) shared-credential listeners only The write on this page runs only against a listener on the shared credential model. A per-user TDS session is read-only and refuses every write, so if your tenant has any database RBAC grant or any row-level security policy — the condition that forces per-user login — this will not run against it. Check which model you are on under [credential models](https://originchaindb.com/docs/compatibility/sqlserver-tds#credentials) first, and write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client) if you are not. what this does Prepares one statement, executes it three times with different values, frees it, and then runs a batch insert inside a transaction. Nothing here is OriginChainDB-specific — it is ordinary JDBC — which is the point. the code ``` // One PreparedStatement, executed many times, then a batch insert. // The driver turns this into the RPC procedure family on its own - // you never name those procedures yourself. try (Connection conn = DriverManager.getConnection(url, user, password)) { try (PreparedStatement ps = conn.prepareStatement( "SELECT id, customer_ref FROM invoices" + " WHERE status = ? AND amount_cents > ?")) { for (String status : List.of("open", "settled", "void")) { ps.setString(1, status); ps.setInt(2, 50000); try (ResultSet rs = ps.executeQuery()) { while (rs.next()) { System.out.println(status + " " + rs.getInt("id")); } } } } // close() frees the server-side handle conn.setAutoCommit(false); try (PreparedStatement ins = conn.prepareStatement( "INSERT INTO invoices (id, customer_ref, amount_cents, status)" + " VALUES (?, ?, ?, ?)")) { for (Object[] row : rows) { ins.setInt(1, (Integer) row[0]); ins.setString(2, (String) row[1]); ins.setInt(3, (Integer) row[2]); ins.setString(4, (String) row[3]); ins.addBatch(); } int[] counts = ins.executeBatch(); // every statement runs System.out.println("inserted " + counts.length + " rows"); } conn.commit(); } ``` what goes over the wire A SQL Server driver does not send that statement the same way twice. It switches to a prepared handle, and which call it uses depends on the driver and on how many times you have executed: ``` execution 1 sp_executesql statement text + parameter declarations execution 2 sp_prepexec prepare and execute in one round trip -> returns a handle execution 3 sp_execute the handle, with fresh values close sp_unprepare free the handle ``` All of it is served. That matters more than it sounds: a listener that answered only the first form would make "the driver works" true for exactly one execution of any prepared statement and false from the second onward — and false from the first for ODBC clients, which reach for the handle immediately. The interop battery asserts the whole family through the driver's own handle, not through our encoder. what this is not Those procedure names belong to the protocol, not to your schema. Serving them is what makes a prepared statement work; it is not stored-procedure support. Calling a procedure you wrote is refused, as is every other procedure by name or by numeric id. ``` -- Calling a procedure YOU wrote is a different thing, and is refused. EXEC dbo.settle_invoice @id = 4101; -- error: that procedure surface does not exist on this engine ``` OriginChainDB's own procedures are written in a PL/pgSQL-shaped subset and called over the [SQL endpoint](https://originchaindb.com/docs/sql) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client). The [TDS reference](https://originchaindb.com/docs/compatibility/sqlserver-tds#procedures) shows the same procedure written both ways. notes - Handles are per connection. They are freed on every teardown path, they are never recycled after a close, and one connection can never see another's. - `executeBatch` runs every statement in the batch. Earlier builds parsed only the first request in a message and dropped the rest, reporting success for statements that never ran; that is fixed and pinned. - OUTPUT parameters are refused, because there is no variable-assignment surface that could produce one. Return values through a result set instead. - Integer, floating point, bit, string, numeric and unique-identifier parameters decode exactly. Date and time parameters are best-effort in this preview — their epoch base differs from the engine's stored form. Microsoft, SQL Server and T-SQL are trademarks of the Microsoft group of companies. Named here only to describe protocol and dialect compatibility. OriginChainDB is not affiliated with, endorsed by, or sponsored by Microsoft, and does not distribute Microsoft software. --- # T-SQL that is refused - OriginChainDB TDS example Canonical source: https://originchaindb.com/docs/examples/sqlserver/tsql-refused Sitemap last modified: 2026-09-13T20:18:50.000Z examples · tds · 5 / 5 # 5. T-SQL that is refused [← SQL Server wire examples](https://originchaindb.com/docs/examples/sqlserver) what this does Everything here comes back as an explicit error naming the construct. Nothing runs, nothing is half-applied, and the connection stays open — a refusal is a normal answer on this surface, not a fault. This page is the honest half of the compatibility story, and the half worth reading first if you are costing a migration. The design rule behind all of it is soundness over coverage: a construct without an exact equivalent is refused rather than approximated, because an approximation returns an answer that looks right. variables and built-in globals ``` -- refused DECLARE @cutoff INT = 100000; SELECT id FROM invoices WHERE amount_cents > @cutoff; SELECT @@VERSION; SELECT @@ROWCOUNT; -- do this instead: bind the value as a parameter, so the driver -- carries it and the engine never sees a variable at all. -- JDBC: ps.setInt(1, 100000) -- Go: db.Query("... amount_cents > @p1", 100000) -- Row counts come back through the driver's own update-count API. ``` Note the difference between a T-SQL variable and a driver placeholder. The Go driver writes `@p1` for a bind parameter and that is fine — it is rewritten to a positional bind before the statement is parsed. A variable you declared yourself is a different thing, and there is nowhere for it to live. temporary objects ``` -- refused SELECT id, amount_cents INTO #recent FROM invoices WHERE status = 'open'; SELECT * FROM #recent; -- do this instead: an ordinary table, or a CTE if the scope is one query. WITH recent AS ( SELECT id, amount_cents FROM invoices WHERE status = 'open' ) SELECT * FROM recent; ``` the GO separator ``` -- refused: GO is a client-tool directive, not a server statement. INSERT INTO invoices (id, status) VALUES (1, 'open'); GO INSERT INTO invoices (id, status) VALUES (2, 'open'); GO -- do this instead: send one statement per batch. Your driver already -- does that unless you are pasting a script written for a query tool. ``` This one is caught before parsing rather than after, because a SQL Server grammar reads a trailing `GO` as a column alias — so `SELECT 1 GO` would parse cleanly and mean something you did not write. A standalone `GO` line is what is rejected; `go` used as a column name is untouched. MERGE ``` -- refused MERGE invoices AS t USING staged AS s ON t.id = s.id WHEN MATCHED THEN UPDATE SET t.status = s.status WHEN NOT MATCHED THEN INSERT (id, status) VALUES (s.id, s.status); -- do this instead INSERT INTO invoices (id, status) SELECT id, status FROM staged ON CONFLICT (id) DO UPDATE SET status = EXCLUDED.status; ``` date and time arithmetic ``` -- refused: there is no calendar timestamp type behind this surface, -- so there is no sound mapping. Timestamps are ISO-8601 strings. SELECT DATEADD(day, -7, GETDATE()); SELECT DATEDIFF(day, created_at, GETDATE()); SELECT CONVERT(VARCHAR(10), created_at, 112); -- style code = formatting -- do this instead: compute the boundary in your application and bind it. -- Instant cutoff = Instant.now().minus(7, ChronoUnit.DAYS); -- ps.setString(1, cutoff.toString()); SELECT id FROM invoices WHERE created_at >= ?; ``` This is the refusal that catches most reporting queries. There is no calendar timestamp type behind this surface to do the arithmetic on, so rather than invent one the translator declines and you do the arithmetic where a real calendar exists — in your application. the rest ``` -- refused: no LIMIT equivalent SELECT TOP 10 PERCENT * FROM invoices; SELECT TOP 10 * FROM invoices ORDER BY amount_cents DESC WITH TIES; -- refused: constrain the statement with a predicate instead DELETE TOP (100) FROM invoices WHERE status = 'void'; -- refused: the INTO clause would be dropped, and the statement would -- report a successful table copy that created nothing. SELECT * INTO invoices_archive FROM invoices; -- do this instead: declare the target, then fill it. CREATE TABLE invoices_archive ( id INT NOT NULL, customer_ref VARCHAR(64), amount_cents INT, status VARCHAR(16) ); INSERT INTO invoices_archive SELECT * FROM invoices; -- refused: ambiguous in T-SQL itself - the column's declared type decides -- whether this concatenates or adds. SELECT 'INV-' + some_column; -- do this instead SELECT CONCAT('INV-', CAST(some_column AS VARCHAR(64))); ``` procedures — the expensive one ``` -- refused: T-SQL procedures are not a surface this engine has. CREATE PROCEDURE dbo.settle_invoice @id INT AS BEGIN UPDATE invoices SET status = 'settled' WHERE id = @id; END; GO EXEC dbo.settle_invoice @id = 4101; -- OriginChainDB's procedures are a PL/pgSQL-shaped subset and are created -- and called over the SQL endpoint or the PostgreSQL wire. This is a -- rewrite in a different language, not a port. ``` If your application's logic lives in stored procedures, this is the line item that decides the size of the project, and no amount of wire compatibility reduces it. Read [the procedures section](https://originchaindb.com/docs/compatibility/sqlserver-tds#procedures) of the TDS reference, which shows one procedure written both ways, before you estimate. Microsoft, SQL Server and T-SQL are trademarks of the Microsoft group of companies. PostgreSQL is a registered trademark of the PostgreSQL Community Association of Canada. Named here only to describe protocol and dialect compatibility. OriginChainDB is not affiliated with, endorsed by, or sponsored by any of them, and does not distribute their software. --- # T-SQL that translates - OriginChainDB TDS example Canonical source: https://originchaindb.com/docs/examples/sqlserver/tsql-translated Sitemap last modified: 2026-09-13T20:18:50.000Z examples · tds · 4 / 5 # 4. T-SQL that translates [← SQL Server wire examples](https://originchaindb.com/docs/examples/sqlserver) shared-credential listeners only The write on this page runs only against a listener on the shared credential model. A per-user TDS session is read-only and refuses every write, so if your tenant has any database RBAC grant or any row-level security policy — the condition that forces per-user login — this will not run against it. Check which model you are on under [credential models](https://originchaindb.com/docs/compatibility/sqlserver-tds#credentials) first, and write through the [HTTP API](https://originchaindb.com/docs/api) or the [PostgreSQL wire](https://originchaindb.com/docs/connect-sql-client) if you are not. what this does A batch that arrives over TDS is parsed with a SQL Server grammar and rewritten into OriginChainDB's SQL before it executes. This page shows the rewrite for each construct that has one, so you can tell at a glance whether a query you already have will run. This is the query language only. The procedural language is not implemented — see [what is refused](https://originchaindb.com/docs/examples/sqlserver/tsql-refused) and the [T-SQL boundary](https://originchaindb.com/docs/compatibility/sqlserver-tds#tsql). bracket identifiers ``` -- you send SELECT [order].[id], [customer name] FROM [order] WHERE [status] = 'open'; -- the engine executes SELECT "order".id, "customer name" FROM "order" WHERE status = 'open'; ``` A plain, non-reserved name loses its brackets. A name with spaces or punctuation, or one that is a reserved word, becomes a double-quoted identifier instead — so `[order]` keeps working. If a bracket identifier ever survives the rewrite, the statement is refused rather than passed on under a name the engine would read differently. paging ``` -- you send SELECT TOP 20 id FROM invoices ORDER BY amount_cents DESC; SELECT id FROM invoices ORDER BY id OFFSET 40 ROWS FETCH NEXT 20 ROWS ONLY; -- the engine executes SELECT id FROM invoices ORDER BY amount_cents DESC LIMIT 20; SELECT id FROM invoices ORDER BY id LIMIT 20 OFFSET 40; ``` `FETCH` is always lowered, never passed through. Left alone the engine would ignore it and return every row — a silently wrong answer, which is exactly what the translator exists to prevent. `PERCENT` and `WITH TIES` have no equivalent and are refused. scalar functions ``` -- you send SELECT ISNULL(status, 'open') AS status, LEN(customer_ref) AS ref_len, CHARINDEX('-', customer_ref) AS dash_at, IIF(amount_cents > 100000, 'L', 'S') AS bucket, CONVERT(NVARCHAR(32), amount_cents) AS amount_text, GETDATE() AS seen_at FROM invoices; -- the engine executes SELECT COALESCE(status, 'open') AS status, LENGTH(customer_ref) AS ref_len, POSITION('-' IN customer_ref) AS dash_at, CASE WHEN amount_cents > 100000 THEN 'L' ELSE 'S' END AS bucket, CAST(amount_cents AS VARCHAR(32)) AS amount_text, now() AS seen_at FROM invoices; ``` `SUBSTRING`, `REPLACE`, `LTRIM`, `RTRIM`, `UPPER` and `LOWER` need no rewrite. Two differences worth knowing before you rely on them: - `LEN` counts trailing spaces; T-SQL trims them. - `GETDATE()` returns an ISO-8601 UTC string, and every call in one statement returns the same value. string concatenation ``` -- translated: both operands are provably string SELECT 'INV-' + customer_ref -> CONCAT('INV-', customer_ref) -- untouched: both operands are numeric, so this stays arithmetic SELECT amount_cents + tax_cents -> amount_cents + tax_cents -- REFUSED: one string, one unknown. In T-SQL the column's declared type -- decides whether this concatenates or adds, and guessing would be wrong -- half the time. Write CONCAT(...) or CAST(...) and say which you meant. SELECT 'INV-' + some_column -> error ``` Where it does translate, note that `CONCAT` treats NULL as empty while T-SQL's `+` propagates it. If NULL handling matters in that expression, write the `CASE` you mean. DDL and type spellings ``` -- you send CREATE TABLE [audit log] ( [id] UNIQUEIDENTIFIER NOT NULL, [note] NVARCHAR(200), [detail] NTEXT ); -- the engine executes CREATE TABLE "audit log" ( id UUID NOT NULL, note VARCHAR(200), detail TEXT ); ``` T-SQL type spellings are normalized in DDL column definitions as well as in casts — but only the ones that change no value representation. The rest are left exactly as you wrote them, on purpose: - Normalized: `NVARCHAR(n)`, `NCHAR`, `NTEXT` and `UNIQUEIDENTIFIER`. - Left alone, and therefore refused by the engine: `BIT` (whose T-SQL literals are `1` and `0`, not `true` and `false`), `MONEY` and `DATETIME2`. Mapping them would silently change what the value means, so you get an error and choose a type yourself. Microsoft, SQL Server and T-SQL are trademarks of the Microsoft group of companies. Named here only to describe dialect compatibility. OriginChainDB is not affiliated with, endorsed by, or sponsored by Microsoft, and does not distribute Microsoft software. --- # Vector search examples - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/vector Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector # Vector examples [← All examples](https://originchaindb.com/docs/examples) Vectors on OriginChainDB live on the runtime endpoint `POST /v1/tenants/:t/vector/:table/put` - they are not declared on the row schema. The `id` field is what links a vector back to a row, so use the row's primary key. 13 examples, each on its own page. The first seven cover the write and the query: insert embeddings, search by cosine, L2 or dot product, switch between fast and high-recall mode, and restrict a search with a metadata filter. The rest cover the parts you cannot guess - deleting vectors, installing and retraining an IVF index, checking what the index actually covers, diagnosing and rebuilding an HNSW graph, building a compressed index, and fusing a sparse and a dense search into one call. Every page carries the request in cURL + Python + TypeScript + Go, the response shape, how it works, and common mistakes. [1 Insert a single embedding Save one vector under an ID. The simplest write. →](https://originchaindb.com/docs/examples/vector/put-single)[2 Insert with metadata Attach arbitrary tags to each vector so you can later filter searches on them. →](https://originchaindb.com/docs/examples/vector/put-metadata)[3 Top-k - cosine similarity The default metric for text embeddings (OpenAI, Cohere, BGE, E5). →](https://originchaindb.com/docs/examples/vector/topk-cosine)[4 Top-k - L2 distance Euclidean distance. Use when the absolute magnitude of vectors carries signal. →](https://originchaindb.com/docs/examples/vector/topk-l2)[5 Top-k - dot product Inner product. Use when your embeddings are already unit-normalised. →](https://originchaindb.com/docs/examples/vector/topk-dot)[6 fast mode - lower latency, slightly lower recall Trade some recall for ~3x faster results. Useful when latency matters more than perfect ranking. →](https://originchaindb.com/docs/examples/vector/topk-fast)[7 Filtered top-k - metadata predicate Restrict the search to vectors whose metadata matches a filter. Applied during search, not after. →](https://originchaindb.com/docs/examples/vector/topk-filtered)[8 Delete a vector Remove one embedding by id, or up to 10,000 in one bulk call. Both are idempotent. →](https://originchaindb.com/docs/examples/vector/delete)[9 Train and install IVF centroids Bootstrap an IVF index: train k-means over the stored corpus, install it, and read back what is installed. →](https://originchaindb.com/docs/examples/vector/centroids)[10 IVF coverage and cell skew How much of the table the index actually covers, how lopsided the cells are, and whether to retrain. →](https://originchaindb.com/docs/examples/vector/ivf-rebalance-status)[11 HNSW health, and rebuilding a damaged index Find out whether the graph can still be walked, prove which metric it was built for, then rebuild it in place. →](https://originchaindb.com/docs/examples/vector/hnsw-health)[12 Build a compressed IVF-PQ index A product-quantized IVF index from a named preset - what it resolves to, and what it keeps on disk. →](https://originchaindb.com/docs/examples/vector/ivf-pq-index)[13 Hybrid top-k - sparse and dense, fused One call that runs both searches and fuses the two rankings server-side. →](https://originchaindb.com/docs/examples/vector/hybrid-topk) --- # Train and install IVF centroids - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/vector/centroids Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 9 / 13 # 9. Train and install IVF centroids [← Vector examples](https://originchaindb.com/docs/examples/vector) ## what this does `POST /v1/tenants/:t/vector/:table/train-and-install-centroids` reads the vectors already stored in the table, runs mini-batch k-means over them, and installs the resulting centroid matrix as the table's IVF partitioning - in one call. It then backfills a cell posting for every stored vector, so a table that was written before the install becomes queryable without a re-ingest. `GET /v1/tenants/:t/vector/:table/centroids` reads back what is installed. ## when to use it - You want to query a table with `"index": "ivf"`. Without installed centroids that query returns 503 - IVF has no partitioning to probe. - The corpus has grown or shifted since the last install and the cells no longer match the data. - You are moving a table off the default HNSW path onto IVF for a corpus large enough that a cell probe beats a graph walk. ## the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/train-and-install-centroids" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "partitions": 1024, "init": "kmeans_plus_plus", "max_iterations": 50, "batch_size": 1024, "convergence_threshold": 1e-5, "seed": 42 }' ``` #### Python ``` result = db.vector.train_and_install_centroids( "shop.products", partitions=1024, init="kmeans_plus_plus", max_iterations=50, seed=42, ) print(result.installed, result.iterations, result.converged) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/train-and-install-centroids`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ partitions: 1024, init: "kmeans_plus_plus", max_iterations: 50, seed: 42, }), }, ); const out = await res.json(); console.log(out.installed, out.populated); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. payload, _ := json.Marshal(map[string]any{ "partitions": 1024, "init": "kmeans_plus_plus", "max_iterations": 50, "seed": 42, }) url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/train-and-install-centroids", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload)) req.Header.Set("Authorization", "Bearer "+token) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` { "trained": true, "installed": true, "partitions": 1024, "dim": 768, "iterations": 37, "converged": true, "last_max_shift": 0.0000073, "training_corpus_size": 100000, "populated": 100000 } // populate_skipped is present only when the backfill did not run. ``` `populated` is the number of stored vectors that got a cell posting from this call. When it is well below `training_corpus_size`, or a `populate_skipped` string is present, the index is installed but not fully covering the corpus - [the coverage report](https://originchaindb.com/docs/examples/vector/ivf-rebalance-status) will show it. ## request fields | Field | Required | Notes | | --- | --- | --- | | partitions | yes | How many Voronoi cells to train. The engine caps this at 65,536, and refuses below the k-means floor described under common mistakes. | | init | no | `"kmeans_plus_plus"` or `"random_sample"`. Defaults to `"random_sample"`. | | max_iterations | no | Mini-batch iteration cap. Defaults to 100. | | batch_size | no | Mini-batch size. Defaults to 1024. | | convergence_threshold | no | Early-stop threshold on the per-iteration maximum centroid shift. Defaults to 1e-4. | | seed | no | PRNG seed. Pass one if you want two runs over the same corpus to produce the same partitioning. | ## reading back what is installed #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/centroids" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` preview = db.vector.centroids("shop.products") print(preview.installed, preview.partitions, preview.dim) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/centroids`, { method: "GET", headers: { Authorization: `Bearer ${OC_TOKEN}`, }, }, ); const out = await res.json(); console.log(out.installed, out.partitions, out.dim); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/centroids", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodGet, url, nil) req.Header.Set("Authorization", "Bearer "+token) resp, err := http.DefaultClient.Do(req) ``` ## what the read-back returns ``` { "installed": true, "partitions": 1024, "dim": 768, "centroids_preview": [ [0.0124, -0.0883, 0.0451, 0.0037, -0.0192, 0.0608, -0.0271, 0.0115] /* ... first 4 centroids, first 8 dims of each ... */ ] } ``` The preview is deliberately truncated to the first 4 centroids and the first 8 dimensions of each, so the response stays small whatever `partitions × dim` is. It is an "is anything installed" check, not a way to export the matrix. A table with nothing installed answers 200 with `"installed": false` and zeroes - not 404. ## how it works The trainer reads up to one million stored vectors, runs mini-batch k-means to `max_iterations` or until the maximum centroid shift falls under `convergence_threshold`, and writes the matrix in a single batch. Steps before that write touch nothing, so a failure leaves the previous partitioning in place. Installing is a full replace, and a replace re-draws the cell boundaries. That is why the install then purges the old cell postings and streams the stored corpus back through the new assignment, one page at a time - reading under a short shared lock, assigning with no lock held, appending under a short exclusive one. The walk is not capped: every stored vector gets a posting. If you already have a centroid matrix from your own training run, `POST /v1/tenants/:t/vector/:table/install-centroids` takes it directly as `{"centroids": [[...], [...]]}` and derives `partitions` and `dim` from the shape of what you send. ## common mistakes - Training on too few vectors. The engine refuses with 400 when the stored count is under `partitions × 4`. Past that floor k-means puts several centroids on the same point. Lower `partitions` or ingest more before training. - Querying IVF before installing. A `"index": "ivf"` query against a table with no centroids returns 503 with `"error": "ivf_centroids_not_installed"` and an `install_url` pointing here. - Assuming it is quick. Training and the backfill both run synchronously inside the request. On a large corpus the call holds the connection for seconds - set a generous client timeout. - Re-installing and expecting the old postings to survive. They do not, and should not: new centroids mean new cells. The purge and re-populate is the point, and re-running the route converges to the same end state. - Reading `centroids_preview` as the matrix. It is four centroids, eight dimensions each, always. --- # Delete a vector - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/delete Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 8 / 13 # 8. Delete a vector [← Vector examples](https://originchaindb.com/docs/examples/vector) ## what this does `DELETE /v1/tenants/:t/vector/:table/:vec_id` removes one embedding. It answers 200 in both cases - the body's `deleted` field carries the outcome, so deleting an id that was never there is a no-op rather than a 404. `POST /v1/tenants/:t/vector/:table/delete-bulk` is the same operation over a list of ids in a single write. ## when to use it - An erasure request - the row is gone from your source of truth and its embedding has to go with it. - Re-embedding under a new model: delete the old vectors before writing the new ones, so a stale vector cannot outrank a fresh one. - Cleaning up after a bad ingest, where you know the ids and want them gone in one call rather than 10,000 requests. ## the request #### cURL ``` curl -X DELETE "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/sku-9281" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` # Idempotent - an id that was never stored comes back deleted=False, not 404. result = db.vector.delete("shop.products", "sku-9281") print(result.deleted) ``` #### TypeScript ``` const result = await db.vectorDelete("shop.products", "sku-9281"); console.log(result.deleted); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/sku-9281", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodDelete, url, nil) req.Header.Set("Authorization", "Bearer "+token) resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` { "deleted": true } ``` `deleted` is `false` when nothing was stored under that id. Nothing is written on that path, so a retry costs you a round trip and nothing else. ## query parameters | Field | Required | Notes | | --- | --- | --- | | index | no | `"hnsw"` (default), `"ivf"` or `"ivf_pq"`. Picks which index family the delete dispatches into. | | repair | no | HNSW only. `true` re-links the deleted node's neighbours across the hole instead of leaving a tombstone. Ignored on the IVF and IVF-PQ arms, which shrink their posting lists cleanly. | ## deleting in bulk #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/delete-bulk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "ids": ["sku-9281", "sku-1144", "sku-5520"], "index": "hnsw", "repair": true }' ``` #### Python ``` result = db.vector.delete_bulk( "shop.products", ids=["sku-9281", "sku-1144", "sku-5520"], repair=True, ) print(result.deleted_count, result.missing_count) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/delete-bulk`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ ids: ["sku-9281", "sku-1144", "sku-5520"], repair: true, }), }, ); const out = await res.json(); console.log(out.deleted_count, out.missing_count); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. payload, _ := json.Marshal(map[string]any{ "ids": []string{"sku-9281", "sku-1144", "sku-5520"}, "repair": true, }) url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/delete-bulk", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload)) req.Header.Set("Authorization", "Bearer "+token) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what bulk delete returns ``` { "deleted_count": 2, "missing_count": 1 } ``` The two counts always sum to the number of distinct ids you sent - duplicates inside one call are de-duplicated server-side and count once. ## how it works On the HNSW arm a delete tombstones the node: the graph slot stays, the embedding goes. `repair: true` additionally re-links the neighbours across the hole, at a higher per-call cost. In bulk that repair runs as a single sweep at the end of the batch rather than once per id. On the IVF and IVF-PQ arms the id is evicted from its cell's posting list and the payload is deleted; there is no topology to repair. Every successful delete also decrements the tenant's stored-embedding counter, so the quota tracks embeddings currently resident rather than embeddings ever written. A bulk delete writes one entry to the destructive-operations audit trail carrying the counts, the index family and whether repair ran. The ids themselves are deliberately not logged - a vector id is customer-chosen and routinely a document or user key. ## common mistakes - Treating a missing id as an error. Both routes return 200 for an id that was not there. Assert on `deleted` / `deleted_count`, not on the status code. - Forgetting `index` on an IVF table. The dispatch defaults to HNSW. Delete an IVF row without `?index=ivf` and the cell posting still lists it. - Sending more than 10,000 ids. That is the per-request cap and it is a 400, not a truncation. Page your own list. - Expecting `repair` to help an IVF delete. It is accepted on both arms so one client shape works everywhere, but it only does anything on HNSW. - Very long ids. A vector id over 1 KiB is refused with 400 - the limit exists so the hash and write path stay bounded. --- # HNSW index health and rebuild - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/vector/hnsw-health Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 11 / 13 # 11. HNSW health, and rebuilding a damaged index [← Vector examples](https://originchaindb.com/docs/examples/vector) ## what this does `GET /v1/tenants/:t/vector/:table/hnsw-health` reports whether a table's HNSW graph can actually be walked: how many live vectors a search can reach, how many disconnected components the graph has, how many nodes nothing points at, and a one-line verdict. `POST /v1/tenants/:t/vector/:table/rebuild-hnsw` rebuilds the graph from the vectors already in the store. No embedding is rewritten and nothing is re-ingested. ## when to use it - Recall is worse than it should be and widening the search beam does not help - that pattern is a graph problem, not a tuning problem. - You inherited a table and do not know which distance metric its index was built for. - Before a fleet-wide reindex, to find which tables actually need one instead of rebuilding everything. ## the request #### cURL ``` curl -G "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/hnsw-health" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "self_recall=512" ``` #### Python ``` # No typed SDK helper for this route yet - call it directly. import httpx r = httpx.get( f"{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.products/hnsw-health", headers={"Authorization": f"Bearer {OC_TOKEN}"}, params={"self_recall": 512}, ) r.raise_for_status() health = r.json() print(health["verdict"]) print(health["self_recall"]["self_recall"]) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/hnsw-health?self_recall=512`, { method: "GET", headers: { Authorization: `Bearer ${OC_TOKEN}`, }, }, ); const health = await res.json(); console.log(health.verdict, health.shattered); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/hnsw-health?self_recall=512", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodGet, url, nil) req.Header.Set("Authorization", "Bearer "+token) resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` { "table": "shop.products", "present": true, "nodes": 41892, "live": 41892, "tombstoned": 0, "reachable_live": 41892, "reachable_fraction": 1.0, "components": 1, "zero_indegree_live": 0, "indegree_percentiles": [1, 4, 9, 16, 27, 48, 91], // min p1 p10 p50 p90 p99 max "mean_outdegree": 15.8, "shattered": false, "damage": [], "build_metric": "cosine", "stored_vectors": 41892, "unindexed_records": 0, "verdict": "healthy: 41892 of 41892 live vectors retrievable, ...", "self_recall": { "metric": "cosine", "applicable": true, "probes": 512, "hits": 512, "self_recall": 1.0, "misses": [], "verdict": "healthy: all 512 probed vectors retrieve themselves ..." } /* plus per-layer shape, norm deciles and advisories */ } ``` `self_recall` is present only when you ask for it. Everything else comes back on every call. ## query parameters | Field | Required | Notes | | --- | --- | --- | | self_recall | no | Run the self-retrieval probe over this many live vectors as well as the topology report. Absent or 0 skips it; the cap is 4096 because each sample is a real query inside this one request. | | metric | no | Metric to probe under. Defaults to the recorded build metric, then the schema's declared distance, then cosine. | | dim | no | Vector width. Only needed when the probe runs and the table has no registered vector dimension. | ## how it works The topology half walks the graph and counts. It catches the failure it was built for - a graph that reaches a small fraction of its own nodes, where no amount of beam widening recovers the rest - and it reports `shattered: true` with the reasons in `damage`. The topology half is structurally blind to one thing: an index built for one distance metric and queried under another. That graph is perfectly connected and reports every reachability field as healthy. `?self_recall=N` asks the other question, and needs no ground truth to do it - under cosine, L2 or Manhattan a vector is its own exact nearest neighbour, so any miss is a defect. Under dot product that is not true, and the probe says so with `"applicable": false` rather than returning a number that would read as a failure. That makes the probe the tool for an unstamped index. Run it once per candidate metric; the one the graph was really built under is the one that comes back at 1.0. Then pass that metric to the rebuild, which records it beside the graph so nothing has to guess again. ## rebuilding the graph #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/rebuild-hnsw" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "metric": "cosine", "dry_run": true }' ``` #### Python ``` # No typed SDK helper for this route yet - call it directly. import httpx r = httpx.post( f"{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.products/rebuild-hnsw", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={"metric": "cosine", "dry_run": True}, timeout=None, ) r.raise_for_status() plan = r.json() print(plan["vectors"], plan["estimated_build_secs"], plan["estimated_peak_bytes"]) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/rebuild-hnsw`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ metric: "cosine", dry_run: true }), }, ); const plan = await res.json(); console.log(plan.vectors, plan.estimated_build_secs); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. payload, _ := json.Marshal(map[string]any{ "metric": "cosine", "dry_run": true, }) url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/rebuild-hnsw", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload)) req.Header.Set("Authorization", "Bearer "+token) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what the rebuild returns ``` { "tenant": "…", "table": "shop.products", "rebuilt": false, // false on a dry run "metric": "cosine", "build_metric_before": null, // null on an index built before the stamp "dim": 768, "vectors": 41892, "estimated_peak_bytes": 137070624, "estimated_build_secs": 55, "health_before": { /* the same report as GET hnsw-health */ }, "health_after": null, // null on a dry run "elapsed_ms": 412 } ``` Drop `dry_run` to run it for real: `rebuilt` becomes `true` and `health_after` carries the post-rebuild report, so one response shows you the before and the after. ## rebuild fields | Field | Required | Notes | | --- | --- | --- | | metric | conditional | Required unless the table has a registered vector schema or a recorded build metric. The engine refuses rather than defaulting - a rebuild under the wrong metric produces a well-connected index for the wrong distance, which no health check can detect. | | dry_run | no | Report the plan - current health, vector count, memory estimate, time estimate - and build nothing. Defaults to false. | | force | no | Rebuild even when the diagnostic reads healthy. Defaults to false so a fleet sweep cannot spend hours reindexing tables that are fine. | | max_bytes | no | Override the per-request memory budget for the build. The default ceiling is 2 GiB; a rebuild whose estimate exceeds it is refused with the estimate in the message. | ## common mistakes - Rebuilding a table under write load. The build holds no lock, but the final swap refuses with 409 if any vector was written or deleted while it ran - overwriting the graph would leave the new embedding on disk and unreachable by every query. Nothing is written on that refusal, so a retry is free; a continuously written table will not converge until you pause writes to it. - Rebuilding a healthy index. That is a 409, on purpose. Add `"force": true` if you mean it. - Guessing the metric. If the table has no schema and no recorded build metric, the rebuild refuses instead of assuming cosine. Use `?self_recall=` to find out which metric the graph matches, then state it. - Reading `self_recall` under dot product. Self-retrieval is not a valid signal there, and the probe reports `"applicable": false` with the numeric fields zeroed. - Calling it on a table with no HNSW index. The rebuild answers 404 - there is no graph to rebuild. An IVF or IVF-PQ table is the usual reason. --- # Hybrid sparse and dense top-k - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/vector/hybrid-topk Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 13 / 13 # 13. Hybrid top-k - sparse and dense, fused [← Vector examples](https://originchaindb.com/docs/examples/vector) ## what this does `POST /v1/tenants/:t/vector/:table/topk_hybrid` runs a dense search and a sparse search over the same table and fuses the two rankings server-side, returning one list. Without it you would issue both searches yourself and merge the results in your own code. The sparse side of the table is written with `POST /v1/tenants/:t/vector/:table/put_sparse`, which takes a vector in the usual sparse layout - parallel index and value arrays plus the full width. ## when to use it - Retrieval for RAG, where a dense embedding finds the paraphrase and a sparse one finds the exact term - a product code, an error string, a surname. - Search where a query is sometimes a sentence and sometimes two keywords, and one ranking alone is wrong for half your traffic. - Anywhere you are already issuing two searches and merging them client-side, paying two round trips for it. ## writing the sparse side #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/put_sparse" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "sku-9281", "indices": [17, 402, 5561], "values": [0.82, 1.40, 0.31], "dim": 30000 }' ``` #### Python ``` # No typed SDK helper for this route yet - call it directly. import httpx r = httpx.post( f"{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.products/put_sparse", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "id": "sku-9281", "indices": [17, 402, 5561], "values": [0.82, 1.40, 0.31], "dim": 30000, }, ) r.raise_for_status() # 201 Created, empty body. ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/put_sparse`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "sku-9281", indices: [17, 402, 5561], values: [0.82, 1.4, 0.31], dim: 30000, }), }, ); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. payload, _ := json.Marshal(map[string]any{ "id": "sku-9281", "indices": []uint32{17, 402, 5561}, "values": []float32{0.82, 1.40, 0.31}, "dim": 30000, }) url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/put_sparse", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload)) req.Header.Set("Authorization", "Bearer "+token) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what the write returns `201 Created` with an empty body, the same as the dense `put`. The id is what ties the two sides together: write the sparse vector under the same id as the dense one and the fusion can see both. ## the hybrid query #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk_hybrid" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "dense_query": [0.0124, -0.0883, /* ... 768 floats ... */], "dense_dim": 768, "dense_metric": "cosine", "sparse_query_indices": [17, 402], "sparse_query_values": [0.9, 1.2], "sparse_dim": 30000, "k": 10, "rrf_k": 60 }' ``` #### Python ``` # No typed SDK helper for this route yet - call it directly. import httpx r = httpx.post( f"{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.products/topk_hybrid", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "dense_query": query_768d, "dense_dim": 768, "dense_metric": "cosine", "sparse_query_indices": [17, 402], "sparse_query_values": [0.9, 1.2], "sparse_dim": 30000, "k": 10, }, ) r.raise_for_status() for hit in r.json(): print(hit["id"], hit["score"]) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/topk_hybrid`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ dense_query: query768d, dense_dim: 768, dense_metric: "cosine", sparse_query_indices: [17, 402], sparse_query_values: [0.9, 1.2], sparse_dim: 30000, k: 10, }), }, ); const hits = await res.json(); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. payload, _ := json.Marshal(map[string]any{ "dense_query": query768d, "dense_dim": 768, "dense_metric": "cosine", "sparse_query_indices": []uint32{17, 402}, "sparse_query_values": []float32{0.9, 1.2}, "sparse_dim": 30000, "k": 10, }) url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/topk_hybrid", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload)) req.Header.Set("Authorization", "Bearer "+token) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` [ { "id": "sku-9281", "score": 0.0325 }, { "id": "sku-1144", "score": 0.0161 }, { "id": "sku-5520", "score": 0.0156 } /* ... up to k entries ... */ ] // A bare array - the same hit shape the single-mode topk routes return. ``` `score` here is the fused rank score, not a distance or a similarity. It is always positive and it compares two hits within one hybrid query - it does not compare to the cosine or L2 scores of an ordinary top-k. ## request fields | Field | Required | Notes | | --- | --- | --- | | dense_query | yes | The dense embedding. Its length must equal dense_dim. | | dense_dim | yes | Dense vector width. Must match the table's dimension. | | dense_metric | no | `"cosine"` (default), `"dot"`, `"l2"` or `"manhattan"`. | | dense_mode | no | `"fast"` or `"high_recall"`, tuning the dense leg only. | | sparse_query_indices | yes | Non-zero component positions. Need not be sorted - the server sorts, and sums the values of any duplicate index. | | sparse_query_values | yes | The values at those positions. Same length as the indices. | | sparse_dim | yes | Full sparse width. Every index must be strictly below it. | | k | yes | How many fused hits to return. | | rrf_k | no | Fusion constant, default 60. Raising it flattens the contribution any single list makes; lowering it sharpens it. | | candidates | no | How many candidates to pull from each list before fusing. Defaults to the larger of 4k and 40, and is clamped to at least k. Larger means better recall and more work per query. | | filter | no | Metadata equality filter, applied to both legs before fusion. Same shape as the filter on the single-mode routes. | ## how it works Each leg produces a ranked list of `candidates` hits. Fusion then scores every id by summing `1 / (rrf_k + rank)` over the lists it appears in, and returns the top `k`. Because the sum is over ranks and not over scores, the two legs need no calibration against each other - which is the reason to fuse this way rather than by adding raw similarities. The consequence worth designing around: a document that places second in both lists beats a document that places first in one and is absent from the other. Appearing in both is the signal. Sparse vectors go into their own inverted index, which walks only the posting lists for the terms your query actually names. A document with no term in common with the query has a zero dot product and no defined rank, so it never surfaces from the sparse leg - it can still reach the result through the dense one. ## common mistakes - Comparing a hybrid score to a top-k score. They are different quantities. A hybrid score near 0.03 is not "worse" than a cosine score of 0.94; it is a rank sum. Compare hybrid scores only to other hybrid scores from the same `rrf_k`. - Writing the sparse and dense sides under different ids. The id is the join. Two ids means two documents, and fusion has nothing to combine. - Sending an index equal to `dim`. Indices are zero-based and must be strictly below the width - an out-of-bounds component is a 400, on the write and on the query. - Sending NaN or an infinity in `values`. Refused at the boundary with 400 rather than allowed to poison a score. - Mismatched array lengths. `indices` and `values` must be the same length, on both the write and the query side. - Calling it as a row-restricted user. The fused score encodes ranks assigned in lists that include rows a row-level policy would hide, so the route refuses for such a caller rather than serving a leak that looks fused. --- # Build an IVF-PQ index - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/ivf-pq-index Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 12 / 13 # 12. Build a compressed IVF-PQ index [← Vector examples](https://originchaindb.com/docs/examples/vector) ## what this does `POST /v1/tenants/:t/vector/:table/create-ivf-pq-index` builds a product-quantized IVF index over the vectors already in the table: it trains the coarse cells and the quantizer together, installs both, and sets the query-time policy - all from one named preset. You pick a profile, not a set of knobs. The engine resolves the cell count from the corpus size and the quantizer width from the vector dimension, and reports back exactly what it chose. ## when to use it - The corpus no longer fits comfortably in memory as raw float vectors and you want the index to carry compressed codes instead. - You are on IVF already and want the smaller footprint that quantized codes buy. - You want a repeatable index build: pass a seed and the same corpus produces the same index. ## the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/create-ivf-pq-index" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "preset": "high_recall", "seed": 42 }' ``` #### Python ``` # No typed SDK helper for this route yet - call it directly. import httpx r = httpx.post( f"{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.products/create-ivf-pq-index", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={"preset": "high_recall", "seed": 42}, timeout=None, ) r.raise_for_status() built = r.json() print(built["partitions"], built["pq_m"], built["keep_raw"]) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/create-ivf-pq-index`, { method: "POST", headers: { Authorization: `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ preset: "high_recall", seed: 42 }), }, ); const built = await res.json(); console.log(built.partitions, built.pq_m); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. payload, _ := json.Marshal(map[string]any{ "preset": "high_recall", "seed": 42, }) url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/create-ivf-pq-index", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload)) req.Header.Set("Authorization", "Bearer "+token) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` { "trained": true, "installed": true, "preset": "high_recall", "partitions": 1024, "pq_m": 64, "pq_bits": 8, "keep_raw": true, "raw_vectors_retained": true, "dim": 768, "training_corpus_size": 100000 } ``` `keep_raw` and `raw_vectors_retained` are different questions and are reported separately on purpose. The first is a query-time policy: may the exact re-rank read the raw vector back. The second is the footprint: are the raw vectors still on disk. ## request fields | Field | Required | Notes | | --- | --- | --- | | preset | yes | `"high_recall"` or `"compressed"`. high_recall uses a finer quantizer and keeps the raw vector so the exact re-rank can run; compressed uses a coarser one and does not re-rank. | | seed | no | PRNG seed for the training run. Defaults to 0. Pass one if you need two builds over the same corpus to match. | | drop_raw_vectors | no | Destructive and opt-in. Discard the raw float embeddings table-wide as part of the build, keeping only the codes. Defaults to false: building an index never destroys data. There is no undo - set it only when the vectors are reproducible from your own source of truth. | ## how it works The build reads the stored corpus first, because the preset resolves against it. The cell count comes from the corpus size - roughly four times its square root, rounded to a power of two and clamped between 64 and 65,536 - and the quantizer's subvector count comes from the vector dimension: the divisor of `dim` (at most 64) whose subspace width lands closest to the preset's target, which is 8 for high_recall and 16 for compressed. Codes are 8 bits wide either way. Then it trains, installs the codebook and the centroids, sets the `keep_raw` policy, and populates the cell postings in bounded chunks so a large corpus never holds the write lock across the whole build. Queries reach the index with `"index": "ivf_pq"` on the ordinary top-k route. On a table built with `keep_raw`, `refine_k_factor` controls how many candidates the exact re-rank re-scores; on a codes-only table there is nothing to re-score against and the field is ignored. ## common mistakes - Asking for high_recall and drop_raw_vectors. That combination is refused with 400, and it should be: high_recall's accuracy comes from the exact re-rank, which reads the raw vector back. Honouring the drop would leave the policy pointing at deleted data and quietly degrade every query. - Rebuilding a table whose raw vectors were dropped. An index build needs float vectors to train on, and codes cannot be turned back into them. That is a 400 telling you to re-ingest - the only fix. - Building on an empty table. 400. Write the vectors first; the preset has nothing to resolve against otherwise. - Reading `keep_raw: false` as "the vectors are gone". It means the exact re-rank will not run. Unless you passed `drop_raw_vectors`, `raw_vectors_retained` is still true and the compressed index costs the raw vectors on top of its codes. - Calling it on a quorum-replicated tenant. The route refuses with 501 there. The build re-writes the whole corpus and is not expressible as one replicated frame, so accepting it could leave a promoted node with a half-built index. --- # IVF rebalance status - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/ivf-rebalance-status Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 10 / 13 # 10. IVF coverage and cell skew [← Vector examples](https://originchaindb.com/docs/examples/vector) ## what this does `GET /v1/tenants/:t/vector/:table/ivf-rebalance-status` reports the shape of an IVF index: how many vectors sit in each cell, how lopsided that distribution is, how many stored vectors have a cell posting at all, and whether the engine thinks you should retrain. It is read-only. Nothing is rebalanced as a side effect - the recommendation is a recommendation. ## when to use it - Your IVF queries return fewer or worse hits than you expect, and you need to know whether the index actually covers the data. - After a bulk ingest, to confirm the new vectors were assigned to cells rather than merely stored. - On a schedule, to catch cell skew building up as the corpus drifts away from the centroids you trained. ## the request #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/ivf-rebalance-status" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` # No typed SDK helper for this route yet - call it directly. import httpx r = httpx.get( f"{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.products/ivf-rebalance-status", headers={"Authorization": f"Bearer {OC_TOKEN}"}, ) r.raise_for_status() status = r.json() print(status["coverage"], status["skew"], status["action"]) ``` #### TypeScript ``` // No typed SDK helper for this route yet - call it directly. const res = await fetch( `${OC_HOST}/v1/tenants/${OC_TENANT}/vector/shop.products/ivf-rebalance-status`, { method: "GET", headers: { Authorization: `Bearer ${OC_TOKEN}`, }, }, ); const status = await res.json(); console.log(status.coverage, status.skew, status.action); ``` #### Go ``` // No typed SDK helper for this route yet - call it directly. url := fmt.Sprintf("%s/v1/tenants/%s/vector/shop.products/ivf-rebalance-status", host, tenant) req, _ := http.NewRequestWithContext(ctx, http.MethodGet, url, nil) req.Header.Set("Authorization", "Bearer "+token) resp, err := http.DefaultClient.Do(req) ``` ## what you get back ``` { "total_live": 100000, "partitions": 1024, "live_per_cell": [112, 96, 143, 87 /* ... one entry per cell ... */], "skew": 1.46, "action": "none", "stored_vectors": 100000, "unindexed": 0, "coverage": 1.0 } ``` `action` is one of `"none"`, `"recommended"` or `"required"`, set from `skew` against the engine's thresholds. ## response fields | Field | Required | Notes | | --- | --- | --- | | total_live | - | Vectors with a cell posting, summed across every cell. | | partitions | - | Installed centroid count. Equals the length of live_per_cell. | | live_per_cell | - | Posting count per cell, indexed by cell id. A cell that was never written reports 0. | | skew | - | `max(live_per_cell) / mean(live_per_cell)`, or 0.0 on an empty corpus. A perfectly even index sits at 1.0. | | action | - | `"recommended"` above a skew of 2.0, `"required"` above 5.0, otherwise `"none"`. | | stored_vectors | - | Vectors stored under this table, counted from the keys, whether or not they are cell-assigned. | | unindexed | - | `stored_vectors - total_live`. Vectors with no cell posting: written before the centroids were installed, or left behind by a skipped backfill. | | coverage | - | `total_live / stored_vectors`, clamped to 1.0. This is the number to alert on. | ## how it works The report walks each cell's posting list and counts entries; it does not decode the payloads behind them. That makes it cheap enough to poll, and it means "live" is a posting count rather than a verified row count - an index with stale postings from an interrupted re-install can over-report `total_live`. Re-training and re-installing is the cleanup, not a per-call scan. `stored_vectors`, `unindexed` and `coverage` exist for one failure shape that every other field hides: a table with ten thousand vectors stored and zero of them indexed. Nothing else on this response raises a flag there - every cell is empty, so `skew` is 0.0 and `action` is `"none"` - while every IVF query against that table returns nothing. Skew matters for latency rather than correctness. A query probes a fixed number of cells; if one cell holds ten times its share, the queries that land on it do ten times the work. That is what the `action` hint is tracking. ## common mistakes - Reading `total_live` alone. It cannot distinguish "the index is small" from "the index is empty and the data is elsewhere". Read `coverage` with it, every time. - Calling it on an HNSW table. There are no centroids, so the route answers 503 with `"error": "ivf_centroids_not_installed"`. That is the same refusal an IVF query gets, and it means the same thing. - Waiting for a rebalance to happen. Nothing is automatic. `"action": "required"` is the engine telling you to call [train-and-install-centroids](https://originchaindb.com/docs/examples/vector/centroids) yourself. - Treating a low `coverage` as a query-tuning problem. No `nprobe` setting reaches a vector that has no posting. Re-install the centroids so the backfill runs. --- # Insert with metadata - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/put-metadata Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 2 / 13 # 2. Insert with metadata [← Vector examples](https://originchaindb.com/docs/examples/vector) what this does A `metadata` object on the put request is stored alongside the embedding and becomes filterable at search time. The call is the same put as [example 1](https://originchaindb.com/docs/examples/vector/put-single) with one extra field. when to use it - You'll later want to search "similar shoes only" - filtering by category, brand, region, etc. - The filter applies during search (not after), so highly-selective filters stay fast. - Store the filterable fields you need. Don't dump the whole row - vector search isn't a record retrieval API. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "sku-9281", "embedding": [0.0124, -0.0883, /* ... 768 floats ... */], "dim": 768, "metric": "cosine", "metadata": { "category": "running-shoes", "price_bucket": "100-200", "in_stock": true } }' ``` #### Python ``` db.vector.put( "shop.products", "sku-9281", embedding_768d, metadata={ "category": "running-shoes", "price_bucket": "100-200", "in_stock": True, }, ) ``` #### TypeScript ``` await db.vectorPut("shop.products", { id: "sku-9281", embedding: embedding768d, dim: 768, metric: "cosine", metadata: { category: "running-shoes", price_bucket: "100-200", in_stock: true, }, }); ``` #### Go ``` err := db.VectorPut(ctx, "shop.products", originchain.VectorPutRequest{ ID: "sku-9281", Embedding: embedding768d, Dim: 768, Metric: "cosine", Metadata: map[string]any{ "category": "running-shoes", "price_bucket": "100-200", "in_stock": true, }, }) ``` filtering at search time The filter is an equality predicate on metadata keys. See [Example 7](https://originchaindb.com/docs/examples/vector/topk-filtered) for the search call. Equality only - for range filters (e.g., `price < 100`), filter the result client-side. common mistakes - Forgetting to bucket continuous values. Filter equality on prices is useless ($129.50 ≠ $129.99). Store a bucket field like `"price_bucket": "100-200"` instead. - Updating metadata. To change a vector's metadata, re-put with the new metadata. Partial metadata updates aren't supported. --- # Insert a single vector - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/put-single Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 1 / 13 # 1. Insert a single embedding [← Vector examples](https://originchaindb.com/docs/examples/vector) what this does POST /v1/tenants/:t/vector/:table/put saves one embedding under an `id`. The first put against a table locks in the `dim` and `metric` - every later put must match. when to use it - You computed one embedding (e.g., from an LLM API) and want to make it searchable. - Use the row's primary key as the vector `id` so search results link back cleanly. - For bulk loads, use the [`/put_bulk`](https://originchaindb.com/docs/vector#put-bulk) endpoint instead - it writes one log frame for the whole batch and takes up to 100,000 vectors per call. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "sku-9281", "embedding": [0.0124, -0.0883, 0.0451, /* ... 768 floats ... */], "dim": 768, "metric": "cosine" }' ``` #### Python ``` db.vector.put( "shop.products", "sku-9281", embedding_768d, # list[float] of length 768 ) ``` #### TypeScript ``` await db.vectorPut("shop.products", { id: "sku-9281", embedding: embedding768d, // number[] of length 768 dim: 768, metric: "cosine", }); ``` #### Go ``` err := db.VectorPut(ctx, "shop.products", originchain.VectorPutRequest{ ID: "sku-9281", Embedding: embedding768d, // []float32 of length 768 Dim: 768, Metric: "cosine", }) ``` what you get back ``` { "ok": true } ``` Re-putting the same `id` overwrites the previous vector. request fields | Field | Required | Notes | | --- | --- | --- | | id | yes | String. Usually the row's primary key. | | embedding | yes | Array of floats. Length must equal `dim`. | | dim | yes | Locked at the first put. | | metric | no | Default `"cosine"`. Locked at the first put. | common mistakes - dim mismatch. If you switch models partway through, all puts after the first model fail with `400 dim_mismatch`. Use a fresh table for the new model. - Float precision. Doubles convert to f32 on the wire. ~7 decimal digits of precision is plenty for similarity but not for exact equality. --- # Top-k cosine similarity - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/topk-cosine Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 3 / 13 # 3. Top-k - cosine similarity [← Vector examples](https://originchaindb.com/docs/examples/vector) what this does POST /v1/tenants/:t/vector/:table/topk takes a query vector and returns the `k` nearest stored vectors ranked by cosine similarity, the default metric. Cosine ignores vector length and only compares direction - this is the right default for text embeddings, where models like OpenAI `text-embedding-3` output direction-bearing vectors. when to use it - Semantic search over text - product descriptions, chat history, docs. - RAG retrieval before an LLM call. - Any embedding model whose output isn't already unit-normalised - cosine handles normalisation for you. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.0124, -0.0883, 0.0451, /* ... 768 floats ... */], "k": 10, "dim": 768, "metric": "cosine" }' ``` #### Python ``` hits = db.vector.topk( "shop.products", query=query_768d, # list[float] of length 768 k=10, metric="cosine", ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, // number[] of length 768 k: 10, dim: 768, metric: "cosine", }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, // []float32 of length 768 K: 10, Dim: 768, Metric: "cosine", }) ``` what you get back ``` { "hits": [ { "id": "sku-9281", "score": 0.9421 }, { "id": "sku-1144", "score": 0.9187 }, { "id": "sku-5520", "score": 0.8903 } /* ... up to k entries ... */ ] } ``` `score` is cosine similarity in `[-1, 1]`. Higher = closer. 1.0 is identical direction, 0 is orthogonal, -1 is opposite. request fields | Field | Required | Notes | | --- | --- | --- | | query | yes | Array of floats. Length must equal the table's locked `dim`. | | k | yes | How many hits to return. Typical: 5 - 50. | | dim | yes | Must match the table's locked dim. | | metric | no | Default `"cosine"`. Must match the metric the table was put with. | common mistakes - Cosine similarity vs cosine distance. Some libraries return `1 - similarity` (a distance, lower = closer). OriginChainDB returns similarity directly - higher = closer. Don't sort the wrong way. - Metric mismatch. If the table was put with `"l2"` you cannot topk it with `"cosine"`. The request returns `400 metric_mismatch`. - Query from a different model. Embeddings from `text-embedding-3` are not comparable to embeddings from `all-MiniLM`. The same dim does not mean the same vector space. --- # Top-k dot product - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/topk-dot Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 5 / 13 # 5. Top-k - dot product [← Vector examples](https://originchaindb.com/docs/examples/vector) what this does Setting `"metric": "dot"` on a topk request ranks by dot product (inner product) - the sum of the element-wise product of query and stored vector, where a higher score is closer. When both sides are unit-normalised (length 1), the dot product equals cosine similarity but skips the per-call normalisation, so it is a bit cheaper. when to use it - You already L2-normalise every vector before storing - common with OpenAI `text-embedding-3-small` and Cohere v3. - You normalise the query the same way before calling topk. - You want the lowest-overhead similarity metric the engine offers. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.0124, -0.0883, 0.0451, /* ... 768 unit-normalised floats ... */], "k": 10, "dim": 768, "metric": "dot" }' ``` #### Python ``` hits = db.vector.topk( "shop.products", query=query_768d, # already L2-normalised to length 1 k=10, metric="dot", ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, // already L2-normalised to length 1 k: 10, dim: 768, metric: "dot", }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, // already L2-normalised to length 1 K: 10, Dim: 768, Metric: "dot", }) ``` what you get back ``` { "hits": [ { "id": "sku-9281", "score": 0.9418 }, { "id": "sku-1144", "score": 0.9180 }, { "id": "sku-5520", "score": 0.8895 } /* ... up to k entries ... */ ] } ``` `score` is the inner product. Higher = closer. If both sides are unit-length the score falls in `[-1, 1]`; if not, the score is unbounded. request fields | Field | Required | Notes | | --- | --- | --- | | query | yes | Array of floats. Normalise it to length 1 before sending. | | k | yes | Number of hits to return. | | dim | yes | Must match the table's locked dim. | | metric | yes | Set to `"dot"`. Must match the metric the table was put with. | common mistakes - Dot on non-normalised vectors. Without normalisation, longer vectors win regardless of direction. A vector that is just "bigger" beats a vector that is actually more relevant. Symptom: the same few IDs always rank first. - Normalising one side and not the other. Normalise both stored vectors and the query, or neither. Mixing the two ruins the ranking. - Switching mid-table. If half the table was put pre-normalised and half wasn't, dot scores are no longer comparable. Re-embed and re-put. --- # Top-k fast mode vs high recall - OriginChainDB Canonical source: https://originchaindb.com/docs/examples/vector/topk-fast Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 6 / 13 # 6. Fast mode [← Vector examples](https://originchaindb.com/docs/examples/vector) what this does Add `"mode": "fast"` to a topk request. The engine searches a narrower part of the index and returns sooner. The top hits are usually the same as `"high_recall"`, but a few correct results near the boundary may be missed - this is the recall/latency trade-off. The default is `"high_recall"`. when to use it - RAG behind a re-ranker. If a cross-encoder re-scores the top 50, a missed item at rank 47 doesn't matter. - Hot dashboards. Suggestion panels, "people also viewed", autocomplete. - Agent inner loops. Many small calls where each ms compounds. when not to use it - First-pass retrieval correctness. If a missed hit at rank 8 means a wrong answer downstream, stay on high recall. - Citation lookup. "Find the source for this sentence" needs the right hit, not a near one. - Small tables. Under ~50k vectors high recall is already fast - fast mode buys you nothing. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.0124, -0.0883, 0.0451, /* ... 768 floats ... */], "k": 10, "dim": 768, "metric": "cosine", "mode": "fast" }' ``` #### Python ``` hits = db.vector.topk( "shop.products", query=query_768d, k=10, metric="cosine", mode="fast", ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, k: 10, dim: 768, metric: "cosine", mode: "fast", }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, K: 10, Dim: 768, Metric: "cosine", Mode: originchain.ModeFast, }) ``` what you get back ``` { "hits": [ { "id": "sku-9281", "score": 0.9421 }, { "id": "sku-1144", "score": 0.9187 } /* ... up to k entries ... */ ] } ``` Same response shape as the default. The scores you see are real - the engine never fabricates a score for a hit it didn't actually compare. the mode field | Value | Behaviour | | --- | --- | | "high_recall" | Default. Wider index sweep. Use when correctness of the top hits matters more than latency. | | "fast" | Narrower sweep. Lower latency, slightly lower recall at the tail of the result list. | common mistakes - Using fast everywhere. It's the right call for some workloads but recall drops noticeably on long-tail queries. Measure end-to-end quality before flipping the switch globally. - Treating fast as an SLA. Fast mode reduces latency, but a hot table or a cold cache still moves the number. Watch p95, not a single sample. --- # Filtered top-k - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/topk-filtered Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 7 / 13 # 7. Filtered top-k [← Vector examples](https://originchaindb.com/docs/examples/vector) what this does Pass a `filter` object to restrict the search to vectors whose stored metadata matches every key. The example below only ranks vectors where `category == "shoes"`. See [put-metadata](https://originchaindb.com/docs/examples/vector/put-metadata) for how to attach the metadata in the first place. The filter is applied during the search, not after - so a selective filter (say, 2% of the table) stays fast and still returns `k` hits. when to use it - Multi-tenant search where `tenant_id` must match. - Faceted search (category, brand, language, region). - Soft-delete: filter on `deleted: false` at query time instead of rebuilding the index. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.0124, -0.0883, 0.0451, /* ... 768 floats ... */], "k": 10, "dim": 768, "metric": "cosine", "filter": { "category": "shoes" } }' ``` #### Python ``` hits = db.vector.topk( "shop.products", query=query_768d, k=10, metric="cosine", filter={"category": "shoes"}, ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, k: 10, dim: 768, metric: "cosine", filter: { category: "shoes" }, }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, K: 10, Dim: 768, Metric: "cosine", Filter: map[string]any{"category": "shoes"}, }) ``` what you get back ``` { "hits": [ { "id": "sku-9281", "score": 0.9421 }, { "id": "sku-1144", "score": 0.9187 } /* ... only rows whose category == "shoes" ... */ ] } ``` Same shape as an unfiltered topk. If fewer than `k` rows match the filter, you get back what exists - no padding. filter rules | Rule | Notes | | --- | --- | | Equality only | No `$gt`, `$in`, range, or regex. Just `key: value` pairs. | | Multiple keys are AND | `{ "category": "shoes", "in_stock": true }` means both must match. | | Exact match, case-sensitive | `"Shoes"` does not match `"shoes"`. Normalise on insert. | | Strings, numbers, booleans | Nested objects and arrays are not filterable. Flatten at put time. | `` `` `common mistakes Range filters. Not supported. If you need price < 100, bucket the price at insert time (price_band: "under_100") and filter on the bucket. Case-mismatched values. Filter values must match exactly. "shoes" and "Shoes" are different. Lowercase on the way in. Filtering on a field that was never stored. The filter doesn't error - it just matches nothing, and you get an empty hits array. Double-check the metadata was set when you called put. ← 6. Fast mode 8. Delete a vector →` --- # Top-k L2 distance - OriginChainDB vector example Canonical source: https://originchaindb.com/docs/examples/vector/topk-l2 Sitemap last modified: 2026-09-13T20:18:50.000Z examples · vector · 4 / 13 # 4. Top-k - L2 distance [← Vector examples](https://originchaindb.com/docs/examples/vector) what this does Setting `"metric": "l2"` on a topk request ranks results by L2 distance - the straight-line (Euclidean) distance between two vectors, where the lowest score is the closest hit. Unlike cosine, L2 cares about magnitude: a long vector and a short one pointing the same way are not close. Use L2 when the length of the vector carries real signal, such as raw image-feature embeddings. when to use it - Image-feature embeddings from un-normalised CNN backbones. - Audio or sensor embeddings where amplitude is meaningful. - Any model whose authors specify L2 as the recommended metric. the request #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.0124, -0.0883, 0.0451, /* ... 768 floats ... */], "k": 10, "dim": 768, "metric": "l2" }' ``` #### Python ``` hits = db.vector.topk( "shop.products", query=query_768d, k=10, metric="l2", ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, k: 10, dim: 768, metric: "l2", }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, K: 10, Dim: 768, Metric: "l2", }) ``` what you get back ``` { "hits": [ { "id": "img-0421", "score": 0.1832 }, { "id": "img-9912", "score": 0.2417 }, { "id": "img-3308", "score": 0.3009 } /* ... up to k entries ... */ ] } ``` `score` is the L2 distance. Lower = closer. 0.0 means identical vectors. This is the reverse of cosine, so don't reuse a "sort descending" assumption from elsewhere in your code. request fields | Field | Required | Notes | | --- | --- | --- | | query | yes | Array of floats. Length must equal the table's locked `dim`. | | k | yes | Number of hits to return. | | dim | yes | Must match the table's locked dim. | | metric | yes | Set to `"l2"`. Must match the metric the table was put with. | common mistakes - Assuming higher score = closer. For L2 the order is reversed. The first hit has the lowest score. If you wrap results in your own UI, double-check the sort direction. - L2 on already-normalised vectors. If every vector has length 1, L2 and cosine give the same ranking but L2 is slightly more expensive. Pick cosine and move on. - Metric mismatch. The metric is locked at table creation. Switching from `"cosine"` to `"l2"` means creating a fresh table. --- # Create an account and a free instance - OriginChainDB Canonical source: https://originchaindb.com/docs/free-account Sitemap last modified: 2026-09-23T08:33:23.000Z Getting started # Create an account and a free instance Verify your account, choose an organization, and prepare your first OriginChainDB instance. Create your accountOriginChainDB console [[Image: Current OriginChainDB signup page with account fields and sign-in navigation.]](https://originchaindb.com/_astro/signup.DUGy_WQv.png) The current account-creation screen. Complete signup before creating a live instance.Current application · Account form ## 1. Create your account Open [Create an account](https://app.originchain.ai/signup) and complete the fields shown in the form. Read the service agreements before accepting them. Submit the form, then verify the six-digit code sent to your email. If the code expires or is rejected, request a new one from the form. A local demo workspace is separate from a verified account. ## 2. Open your organization After signing in, use the organization selector to choose the workspace that should own the database. Open Instances and choose New instance. ## 3. Choose the free configuration Select Free, enter the instance details, and review the provider and region shown by the console. Check the live plan's storage quota, eligibility, and idle behavior before confirming. Configure an instanceOriginChainDB console [[Image: Current instance creation page showing the configuration step and local preview context before review.]](https://originchaindb.com/_astro/launch.BWiyIGxh.png) The current instance-creation flow, shown with local preview data.Current console · Dark theme · Local preview The captured console is a local preview. Its 500 MB sample configuration and Create preview instance button do not provision a live database or issue credentials. Follow the [instance-creation guide](https://originchaindb.com/docs/dashboard/create-instance) for that distinction. ## 4. Confirm the instance is ready For a live launch, wait for the control plane to confirm that provisioning has completed. Inspect the instance's status and connection details. If the free service sleeps while idle, allow its next request to wake it; use the live status rather than a sample preview badge. ## 5. Connect your application Open Connect and choose an application API or database-client guide. Use the endpoint and credentials issued for your instance, not the sample placeholders shown in a preview. Open the connection guideOriginChainDB console [[Image: Connection guide open over the current instance overview, with the application header and sidebar visible.]](https://originchaindb.com/_astro/connect.DDybhNZM.png) The connection guide organizes client instructions and connection details in one place.Current console · Dark theme · Local preview If a token is displayed only once, save it securely. The server stores a token hash, so losing the value requires rotation; replacing it invalidates the previous token. Next, follow the [quickstart](https://originchaindb.com/docs/quickstart) to create a table and run your first query. --- # Full-text search on OriginChainDB - BM25, boolean, phrase Canonical source: https://originchaindb.com/docs/fts Sitemap last modified: 2026-09-23T08:33:23.000Z reference · full-text # Full-text search Find matching documents with keyword, phrase, and boolean queries, or rank results with BM25. Full-text indexes live on their own runtime endpoint - they are not declared on the schema. You index a text under `(table, field, doc_id)`, then search. The `doc_id` is what links search hits back to your rows - use the row's primary key. `:table` and `:field` are opaque path segments - they are not validated against any schema. They must match exactly between the index call and the query call. Index under `shop.products` and query `shop_products` and you get a silent empty result, not an error. The default search mode is `boolean` - omit `mode` and you get AND-of-terms. ## 1. Index a text. what this does Tell OriginChainDB to make a piece of text searchable. Re-indexing the same `doc_id` replaces the old text in the same write - no stale matches. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "sku-9281", "text": "Lightweight road runner with a carbon plate, designed for marathon pace." }' ``` #### Python ``` db.fts.index( "shop.products", "description", doc_id="sku-9281", text="Lightweight road runner with a carbon plate, designed for marathon pace.", ) ``` #### TypeScript ``` await db.ftsIndex("shop.products", "description", { doc_id: "sku-9281", text: "Lightweight road runner with a carbon plate, designed for marathon pace.", }); ``` #### Go ``` err := db.FTSIndex(ctx, "shop.products", "description", originchain.FTSIndexRequest{ DocID: "sku-9281", Text: "Lightweight road runner with a carbon plate, designed for marathon pace.", }) ``` what you get back ``` HTTP/1.1 201 Created (empty body) ``` A successful index returns `201 Created` with no body. There is no token count or confirmation JSON - check the status code, not the body. what each field means | Field | Where | What it is | | --- | --- | --- | | :table | URL | An opaque label for the index - conventionally your row table, e.g. `shop.products`. Not checked against any schema; must match byte-for-byte at query time. | | :field | URL | Which "logical column" you're indexing under. You can index multiple fields per table - `title`, `description`, etc. | | doc_id | body | A unique ID for this document. Use the row's primary key so hits link back cleanly. | | text | body | The text to search. No size limit on this endpoint, but very large documents are better split into multiple `doc_id`s. | common mistakes - Indexing only one of several fields. If you want users to search "Wireless headphones" and match products whose title or description contains those words, you need to either concatenate both fields into one text before indexing, or index each field separately and union the results. - Forgetting to re-index on updates. Editing a row's text does not automatically update the FTS index. Re-call this endpoint with the new text whenever you change the source. ## 2. BM25 - ranked search. what this does Return the top-k documents ranked by relevance to the query. This is what most users mean when they say "search". Rare query words count more than common ones; documents where the query words appear more often (relative to length) rank higher. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description?q=carbon+marathon&mode=bm25&k=10" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` res = db.fts.search( "shop.products", "description", "carbon marathon", mode="bm25", k=10, ) for hit in res.hits: print(hit.score, hit.doc_id) ``` #### TypeScript ``` const hits = await db.ftsSearch("shop.products", "description", { q: "carbon marathon", mode: "bm25", k: 10, }); for (const hit of hits) { if (typeof hit === "object") console.log(hit.score, hit.doc_id); } ``` #### Go ``` hits, err := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "carbon marathon", Mode: "bm25", K: 10, }) if err != nil { /* handle */ } for _, h := range hits { fmt.Println(h.Score, h.DocID) } ``` what you get back ``` [ { "doc_id": "sku-9281", "score": 9.42 }, { "doc_id": "sku-3140", "score": 4.18 } ] ``` Plain BM25 returns a bare array of `{ doc_id, score }`, best first - no `{ "mode": ..., "hits": [...] }` wrapper. The wrapper object only appears when you add `highlight=true` or `facets=` (then you get `{ "hits": [...], "facets": {...} }`); `explain=true` returns a separate scoring-breakdown object. query params — all BM25-only; ignored in boolean / phrase | Param | Required | Notes | | --- | --- | --- | | q | yes | The query text. URL-encode spaces as `+` or `%20`. | | mode | no | `bm25` \| `boolean` \| `phrase`. Defaults to `boolean` when omitted. Use `bm25` for this section. | | k | no | Max results. Default 10. BM25 only. | | fuzzy | no | Edit distance for typo tolerance. `fuzzy=1` matches one-character typos. | | highlight | no | `highlight=true` returns matched-term snippets per hit. | | facets | no | Comma-separated field names to aggregate as facet counts. Switches the response to the wrapper-object shape. | | explain | no | `explain=true` returns a BM25 scoring-breakdown object instead of hits. | `highlight=true` requires the doc text to have been stored first via `POST /fts/:t/:f/doc`. All of `k`, `fuzzy`, `highlight`, `facets`, and `explain` are silently ignored in boolean and phrase modes. ## 3. Boolean AND - every word must match. what this does Return every document that contains all the query words, in any order, with no ranking. Fast token-presence check - use when you don't need relevance scoring. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description?q=carbon+marathon&mode=boolean" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` hits = db.fts.search( "shop.products", "description", q="carbon marathon", mode="boolean", # every word must appear ) for hit in hits: print(hit.doc_id) ``` #### TypeScript ``` const hits = await db.ftsSearch("shop.products", "description", { q: "carbon marathon", mode: "boolean", }); ``` #### Go ``` hits, err := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "carbon marathon", Mode: "boolean", }) ``` what you get back ``` ["sku-9281", "sku-3140"] ``` A bare array of `doc_id` strings (sorted lexicographically) - no `score` field and no wrapper object. If you need ordering by relevance, use BM25. ## 4. Phrase - exact word order. what this does Match documents that contain the query words contiguously, in the exact order given. Use for branded phrases ("New York Times"), product model numbers, log message templates. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description?q=carbon+plate&mode=phrase" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` hits = db.fts.search( "shop.products", "description", q="carbon plate", mode="phrase", # exact phrase "carbon plate", in that order ) ``` #### TypeScript ``` const hits = await db.ftsSearch("shop.products", "description", { q: "carbon plate", mode: "phrase", }); ``` #### Go ``` hits, err := db.FTSSearch(ctx, "shop.products", "description", originchain.FTSSearchRequest{ Q: "carbon plate", Mode: "phrase", }) ``` what you get back ``` ["sku-9281"] ``` Same shape as boolean - a bare array of `doc_id` strings. Phrase just narrows which docs qualify. ## 5. End-to-end - index, search, enrich. A complete runnable sequence against a live tenant. Set `$OC_HOST`, `$OC_TENANT`, and `$OC_TOKEN` first. Every response shown below is the real shape the engine returns. Steps 3 and 4 use `mode=bm25`, which needs the Full-Text Pro capability enabled on the instance - it is included on every paid configuration at no extra charge. Indexing and the boolean query in step 2 need nothing. step 1 — index three docs ``` # 1) Index three docs. Each POST returns 201 with an EMPTY body. curl -sS -o /dev/null -w "%{http_code}\n" \ -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{ "doc_id": "p001", "text": "Wireless over-ear headphones with active noise cancellation" }' # → 201 curl -sS -o /dev/null -w "%{http_code}\n" \ -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{ "doc_id": "p002", "text": "Wired earbuds, no noise cancellation" }' # → 201 curl -sS -o /dev/null -w "%{http_code}\n" \ -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{ "doc_id": "p003", "text": "USB-C charging cable, 2 metres" }' # → 201 ``` step 2 — boolean query (default mode) ``` # 2) Boolean query (the DEFAULT mode). Bare array of doc_id strings. curl -sS -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=wireless noise" # → ["p001"] (only p001 has BOTH "wireless" AND "noise") ``` step 3 — bm25 ranked query ``` # 3) BM25 ranked query. Bare array of { doc_id, score }, best first. curl -sS -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=noise cancellation" \ --data-urlencode "mode=bm25" \ --data-urlencode "k=10" # → [ { "doc_id": "p001", "score": 6.31 }, { "doc_id": "p002", "score": 2.04 } ] ``` step 4 — enrich with highlights ``` # 4) Enrich with highlights. First store the doc text (highlights read it), # THEN ask for highlight=true. Now the response is an OBJECT, not an array. curl -sS -o /dev/null -w "%{http_code}\n" \ -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description/doc" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{ "doc_id": "p001", "text": "Wireless over-ear headphones with active noise cancellation" }' # → 201 curl -sS -G "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ --data-urlencode "q=noise cancellation" \ --data-urlencode "mode=bm25" \ --data-urlencode "highlight=true" # → { # "hits": [ # { "doc_id": "p001", "score": 6.31, # "highlights": { "description": ["…active noise cancellation"] } } # ] # } ``` ## 6. Analyzer + languages. read this first Stemming, lemmatization, diacritic folding, and stopword removal are implemented in the engine but not yet selectable through the HTTP API. There is no language or analyzer query parameter today. The analyzer the API actually uses is Unicode tokenize + lowercase, and nothing else. Concretely: `q=runs` will not match a document that says "running", and `q=cafe` will not match "café". Plan your indexing around exact tokens. The one transform that is live is synonyms (see below). ### What runs today - Unicode tokenize. Text is split into words by the Unicode word-boundary rules (UAX #29), so it works across scripts. - Lowercase. Every token is lowercased, so `Wireless` and `wireless` match. This is the only normalisation applied. - Synonyms. If you install a per-(table, field) synonym map via `POST /fts/:t/:f/synonyms`, members of a class are treated as equivalent at both index time and BM25 query time. This is the one customer-controlled analyzer feature that is wired through. See [FTS runtime calls](https://originchaindb.com/docs/schemas/fts). ### Built in the engine, not yet exposed (roadmap) The following analyzer stages exist in the engine but cannot be turned on from the API yet. They are listed so you know what is coming, not what you can call today. | Step | What it would do | Status | | --- | --- | --- | | fold diacritics | "café" would match "cafe". | not API-exposed | | stopwords | Drop common words ("the", "and", "of"). | not API-exposed | | stemming | Suffix-strip. "running" / "runs" → "run". Snowball-based. | not API-exposed | | lemmatization | Dictionary lookup. "ran" → "run". More precise than stemming. | not API-exposed | languages the engine has stemmers for (not yet selectable) ArabicDanishDutchEnglishFinnishFrenchGermanHungarianItalianNorwegianPortugueseRomanianRussianSpanishSwedishTamilTurkishHindi There is no per-field analyzer knob you can set today. The runtime calls that are live - plain / JSON-aware index, doc store, synonyms, stopword override - are documented on this page and in [FTS runtime calls](https://originchaindb.com/docs/schemas/fts). ## 7. Examples. 7.1 ## Search — a single term. Search is a `GET` on the same path you indexed to, with the query in `?q=`. `mode=bm25` is what you want when you care about ranking. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=delivery&mode=bm25&k=10" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` res = db.fts.search("shop.orders", "notes", "delivery", mode="bm25", k=10) for hit in res.hits: print(hit.doc_id, hit.score) ``` #### TypeScript ``` const hits = await db.ftsSearch("shop.orders", "notes", { q: "delivery", mode: "bm25", k: 10, }); // mode "bm25" returns RankedHit[]; "boolean" and "phrase" return string[]. for (const h of hits as { doc_id: string; score: number }[]) { console.log(h.doc_id, h.score); } ``` #### Go ``` hits, err := db.FTSSearch(ctx, "shop.orders", "notes", originchain.FTSSearchRequest{ Q: "delivery", Mode: "bm25", K: 10, }) for _, h := range hits { fmt.Println(h.DocID, h.Score) } ``` response — mode=bm25 ``` [ { "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "score": 9.4213 }, { "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB9", "score": 6.1077 } ] ``` A bare array of `{ doc_id, score }`, highest score first. `k` caps the list and defaults to 10. the response shape changes with the mode This endpoint returns four different shapes depending on the parameters. `boolean` and `phrase` give you a bare array of doc_id strings; `bm25` gives an array of objects; adding `highlight` or `facets` wraps it all in an object with a `hits` key; and `explain=true` returns a scoring report instead. The TypeScript SDK models this as a union you have to narrow yourself. 7.2 ## What the query syntax really is. This is worth being blunt about, because it is easy to assume otherwise: `?q=` is not a query language. There is no parser. The string is tokenized into words — Unicode word segmentation, then lowercased — and every token becomes a term. All punctuation is discarded. The practical consequence is that operators you might type do not work, and fail silently rather than erroring. `rush AND delivery` searches for three terms — `rush`, `and`, `delivery`. Quotes around a phrase are dropped. | You might try | What actually happens | | --- | --- | | rush AND delivery | Searches `rush`, `and`, `delivery`. Use `mode=boolean`, which already ANDs. | | rush OR delivery | Searches three terms. Use `mode=bm25`, which already ORs. | | NOT cancelled / -cancelled | Negation is not available here at all. Use `must_not` in the [DSL](https://originchaindb.com/docs/fts#dsl). | | "signed by recipient" | Quotes are discarded. Use `mode=phrase`. | | deliv* | The `*` is discarded. Prefix and wildcard queries exist only in the [DSL](https://originchaindb.com/docs/fts#dsl). | | delivary~1 | This one works. `~N` is the single inline operator the query surface honours. See [fuzzy](https://originchaindb.com/docs/fts#fuzzy). | ### How multiple terms combine — it depends on the mode. This is the important distinction, and it is not configurable: | mode | Multi-term meaning | Returns | | --- | --- | --- | | boolean | AND — every term must be present | Unranked `doc_id` array, lexicographic. Unbounded — `k` is ignored. | | bm25 | OR — any term matches, more/rarer terms score higher | Ranked `{doc_id, score}`, capped at `k`. | | phrase | Adjacent and in order | Unranked `doc_id` array. `k` is ignored. | `` `boolean — every term required cURL Python TypeScript Go mode=boolean (the default) # mode=boolean is the DEFAULT when ?mode= is omitted. # Every term must be present. Returns a bare array of doc_ids, unranked. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=rush+delivery" \ -H "Authorization: Bearer $OC_TOKEN"# The typed namespace defaults to bm25 - pass mode explicitly for boolean. res = db.fts.search("shop.orders", "notes", "rush delivery", mode="boolean") # The legacy method defaults to boolean already: doc_ids = db.fts_search("shop.orders", "notes", q="rush delivery")const docIds = await db.ftsSearch("shop.orders", "notes", { q: "rush delivery", mode: "boolean", }) as string[];docIDs, err := db.FTSSearch(ctx, "shop.orders", "notes", originchain.FTSSearchRequest{ Q: "rush delivery", Mode: "boolean", }) response ["01JTRX9KQ3YH8K2WMX0F5JZAB7", "01JTRX9KQ3YH8K2WMX0F5JZAC1"] an unknown mode falls back to boolean, silently ?mode=ranked, ?mode=BM25 or any typo is not an error — it takes the default branch and runs a boolean AND. If you get back an array of bare strings when you expected scores, check the spelling of mode first. The accepted values are lowercase boolean, bm25 and phrase. the dashboard's search box is different The console workbench accepts a one-line table:field terms shorthand in Search mode. That colon syntax is parsed by the console, which then calls the endpoint documented here — it is not something the engine understands. Don't put shop.orders:notes into ?q=. See Run queries from the dashboard.` 7.3 ## Phrase queries. `mode=phrase` requires the terms to appear adjacent and in the order given. The quotes you would type in another search engine are not syntax here — the mode is the switch. #### cURL ``` # Terms must appear adjacent, in this order. Quotes are NOT syntax - # they would simply be discarded by the tokenizer. Use mode=phrase. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=signed+by+recipient&mode=phrase" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` doc_ids = db.fts.search( "shop.orders", "notes", "signed by recipient", mode="phrase", ) ``` #### TypeScript ``` const docIds = await db.ftsSearch("shop.orders", "notes", { q: "signed by recipient", mode: "phrase", }) as string[]; ``` #### Go ``` docIDs, err := db.FTSSearch(ctx, "shop.orders", "notes", originchain.FTSSearchRequest{ Q: "signed by recipient", Mode: "phrase", }) ``` Phrase results come back as an unranked array of `doc_id` strings — position matching selects the documents, but this mode does not score them. If you need ranking as well as adjacency, use the `phrase` clause inside the [DSL](https://originchaindb.com/docs/fts#dsl), which selects on positions and then ranks with BM25. 7.4 ## Fuzzy matching. Two ways in, both `bm25`-only: a whole-query budget via `?fuzzy=N`, or a per-term `~N` suffix. Putting a `~` anywhere in `q` switches the query onto the fuzzy path automatically. #### cURL ``` # Whole-query budget via ?fuzzy= (bm25 only, 0-3) curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=delivary&mode=bm25&fuzzy=1" \ -H "Authorization: Bearer $OC_TOKEN" # Or per-term inline with ~N. A bare ~ means distance 2. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=delivary~1+recipient&mode=bm25" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` res = db.fts.search("shop.orders", "notes", "delivary", mode="bm25", fuzzy=1) ``` #### TypeScript ``` // The TypeScript SDK does not expose fuzzy - call the endpoint directly, // or embed the ~N operator in the query string. const hits = await db.ftsSearch("shop.orders", "notes", { q: "delivary~1", mode: "bm25", }); ``` #### Go ``` // The Go SDK does not expose fuzzy - embed the ~N operator instead. hits, err := db.FTSSearch(ctx, "shop.orders", "notes", originchain.FTSSearchRequest{ Q: "delivary~1", Mode: "bm25", }) ``` - Edit distance is capped at 3. Higher is a `400`: `fuzzy edit_distance 5 exceeds MAX_EDIT_DISTANCE 3`. - A bare `term~` with no number means distance 2. `term~0` is an exact match. - Each term expands to at most 50 dictionary candidates. - Distance 1 catches most real typos. Distance 2 and 3 expand aggressively and will pull in unrelated words — measure before shipping either. - Only the Python SDK exposes a `fuzzy` parameter. In TypeScript and Go, put `~N` in the query string. 7.5 ## Ranking, scoring and explain. `mode=bm25` scores each document as the sum of its query terms' contributions. Each contribution combines three things: - Inverse document frequency. A term in few documents is worth more than one in many. This is why a rare word dominates a query. - Term frequency, with diminishing returns. The fifth occurrence adds much less than the second. `k1` controls how fast that saturates. - Length normalisation. A hit in a short document counts for more than the same hit buried in a long one. `b` controls how strongly. Defaults are `k1 = 1.2` and `b = 0.75` — the standard Lucene values. On the `?q=` surface they are fixed; they can only be overridden through the [DSL](https://originchaindb.com/docs/fts#dsl)'s `params` object. Scores are relative within one result set — never compare a score across two different queries. ### The explain parameter. Add `explain=true` to a `bm25` query to get the full arithmetic instead of the hits. This is a query parameter on the search route — there is no separate explain endpoint. ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=rush+delivery&mode=bm25&k=5&explain=true" \ -H "Authorization: Bearer $OC_TOKEN" ``` response ``` { "query_terms": ["rush", "delivery"], "n_total": 1204, "avgdl": 11.6, "k1": 1.2, "b": 0.75, "hits": [ { "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "score": 9.4213, "terms": [ { "term": "rush", "df": 42, "idf": 3.361, "tf": 1.0, "doc_len": 5, "contribution": 5.9102 }, { "term": "delivery", "df": 310, "idf": 1.362, "tf": 1.0, "doc_len": 5, "contribution": 3.5111 } ] } ] } ``` `n_total` is the corpus size and `avgdl` the average document length — the two corpus statistics the formula needs. Per hit, `terms` is sorted by `contribution` descending, so the first entry is the term that actually won the document its place. Length normalisation is folded into `contribution`; `doc_len` is the raw token count. - Explain works only with `mode=bm25`, and overrides `highlight` and `facets` if you pass them together. - It always reports the default `k1` and `b` — you cannot explain a custom-tuned score, and the DSL endpoint has no explain of its own. - Explain takes the exhaustive scoring path, so it is slower than the equivalent ranked query. It is a debugging tool, not a production one. - No SDK exposes explain. Call the endpoint directly. 7.6 ## Highlights and facets. Both features need the document's text stored, which the plain index call does not do. Send it to the `/doc` variant instead, along with any facet values you want to aggregate on. ``` # Store the doc text plus per-facet values. Required before highlight=true # or facets= will return anything. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes/doc" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "text": "rush delivery, signed by recipient", "facets": { "status": ["paid"], "channel": ["web"] } }' ``` Then ask for them at query time: ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=rush&mode=bm25&k=5&highlight=true&facets=status,channel" \ -H "Authorization: Bearer $OC_TOKEN" ``` response — note the object wrapper ``` { "hits": [ { "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "score": 9.4213, "highlights": { "notes": ["rush delivery, signed by recipient"] } } ], "facets": { "status": [ { "value": "paid", "count": 12 }, { "value": "refunded", "count": 3 } ], "channel": [ { "value": "web", "count": 11 }, { "value": "app", "count": 4 } ] } } ``` python ``` res = db.fts.search( "shop.orders", "notes", "rush", mode="bm25", k=5, highlight=True, facets=["status", "channel"], ) for hit in res.hits: print(hit.doc_id, hit.score, hit.highlights) for value, bucket in res.facets.items(): print(value, [(b.value, b.count) for b in bucket]) ``` - Highlights come back as raw `` markup around matched terms. Escape or sanitise before rendering. - Facets aggregate; they do not filter. `facets=status` tells you the distribution across the hit set — it does not narrow it. For real filtering use the [DSL](https://originchaindb.com/docs/fts#dsl). - At most 1000 distinct values are tracked per facet field. - Only the Python SDK wraps `highlight` and `facets`; the `/doc` write is not wrapped by any SDK. 7.7 ## Filters and multi-field — the JSON DSL. Everything the `?q=` surface can't do — boolean composition, negation, filters, prefix and wildcard matching, multi-field search with weights, custom `k1`/`b` — lives in a JSON query DSL at `POST /fts/:table/_search`. Note the path takes a table only; fields are named inside the query. ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/_search" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_values_field": "notes", "query": { "bool": { "must": [ { "match": { "field": "notes", "query": "rush delivery" } } ], "should": [ { "term": { "field": "notes", "value": "signed", "boost": 2.0 } } ], "must_not": [ { "term": { "field": "notes", "value": "cancelled" } } ], "filter": [ { "numeric_range": { "field": "amount_cents", "gte": 10000 } } ] } }, "top_k": 20, "params": { "k1": 1.2, "b": 0.75 } }' ``` response ``` { "total": 431, "hits": [ { "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "score": 14.8802 }, { "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAC1", "score": 11.2044 } ] } ``` `total` is the match count before `top_k` truncation, so you can render "showing 20 of 431" without a second query. `top_k` defaults to 10. fields are a named key, not a dynamic one If you know Elasticsearch, this is the one difference that will trip you up. Where ES writes `{"term": {"notes": "rush"}}`, this DSL writes `{"term": {"field": "notes", "value": "rush"}}`. The field is always an explicit `field` key. ### Clause types. `match_all`, `match_none`, `term`, `terms`, `match`, `phrase`, `prefix`, `wildcard`, `regex`, `range`, `numeric_range`, `fuzzy`, `exists`, `multi_match`, `bool`, `constant_score` and `function_score`. - `bool` scores `must` plus any matching `should`. `filter` gates without contributing score; `must_not` excludes. - `match` defaults to OR across its terms; pass `"operator": "and"` to require them all. - Every clause takes a `boost`, and boosts multiply down the tree — a `boost: 2.0` clause inside a `boost: 3.0` bool contributes 6×. - `minimum_should_match` accepts an integer, a negative integer ("all but N"), or a percentage string like `"75%"`. - `regex` supports literals, character classes, repetition, alternation and grouping — but not back-references, look-around or named groups. ### Multi-field search with weights. ``` # Search several indexed fields at once, with per-field weights. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/_search" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": { "multi_match": { "fields": [ { "field": "notes", "boost": 2.0 }, { "field": "customer", "boost": 1.0 } ], "query": "rush delivery", "type": "best_fields", "tie_breaker": 0.3 } }, "top_k": 20 }' ``` `best_fields` (the default) takes the best-scoring field plus `tie_breaker` × each other match; `most_fields` sums them all. `tie_breaker` defaults to `0.0` — pure winner-takes-all — and must be within `[0, 1]`. Every field you name must be independently indexed, or the request 404s. numeric filters need doc_values_field `numeric_range`, `field_value_factor` and decay functions read per-document values that only exist if you wrote them as facets through the `/doc` endpoint. You must also name the source with a top-level `doc_values_field`. Omit it and the request is refused with a 400 rather than quietly matching nothing — a deliberate choice, since a silent empty result is indistinguishable from "no matches". no SQL integration Full-text search is not reachable from `POST /sql`. There is no `MATCH()` function and no way to put a text predicate in a `WHERE` clause. To combine the two, search first and then query the returned `doc_id`s with SQL. Two sibling endpoints share the DSL: `POST /fts/:table/_aggs` for aggregations and `POST /fts/:table/_suggest` for prefix suggestions. Because of these routes, `_search`, `_aggs` and `_suggest` are reserved and cannot be used as field names. No SDK wraps the DSL endpoints. 7.8 ## Limits and gotchas. | Limit | Value | | --- | --- | | Max `k` / `top_k` | 10,000 | | Default `k` / `top_k` | 10 | | Max fuzzy edit distance | 3 | | Fuzzy expansions per term | 50 | | DSL query nesting depth | 32 | | DSL clauses per query | 1024 | | Synonyms per term | 32 | | Distinct values per facet field | 1000 | | Max query string length | no limit | | Pagination / offset | not supported | there is no pagination No `offset`, no `from`, no cursor. You get the top `k` and that is all. To show page two, raise `k` and slice client-side — and remember `boolean` and `phrase` mode ignore `k` entirely and return every match, which is why a broad boolean query on a large corpus can trip the result-size cap and return `413`. top_k bounds the response, not the work A small `k` does not make a broad query cheap. The engine scores every matching document and ranks afterwards, so `match_all` or a very common single term over a large corpus is expensive however few results you ask for. A block-max optimisation skips provably-losing blocks for ranked queries, but it declines to engage in several cases — including whenever you request explain — and falls back to exhaustive scoring. an unindexed field is an empty result on ?q=, a 404 on the DSL The two surfaces disagree here. `_search` refuses an unindexed field with a `404` explaining that an unindexed field is refused "rather than answered with an empty result you could not tell from 'nothing matched'". The `?q=` route has no such check and returns an empty array, so a typo in the path looks exactly like a genuine miss. honest scale guidance Measured on the published benchmark: 50,000 documents → 587 MB on disk, ranked p99 1.8 ms. 200,000 documents → 2.3 GB on disk, p99 13 ms. Cost is roughly 12 KB per document, and resident memory grows linearly with the corpus — projected around 25 GB at a million documents. Latency is production-grade at these sizes; memory, not speed, is what will bound you. Size the instance's RAM against your document count. what needs the capability enabled Indexing documents and `mode=boolean` searches work on any instance. Ranked `bm25` and `phrase` queries, and all three DSL endpoints, need Full-Text Pro enabled on the instance. It is included on every paid configuration at no extra charge - there is nothing to buy - but it is still a capability the instance carries, so until it is switched on the call returns `402` naming it. Turn it on from Billing → Add-ons. caps refuse rather than truncate When a query exceeds an expansion or clause limit the engine returns `400` with a message naming the limit and suggesting a fix — it never silently returns a partial result set. A truncated expansion would drop matching documents without telling you, so the refusal is deliberate. --- # Full-text lemmatization on OriginChainDB Canonical source: https://originchaindb.com/docs/fts/lemmatization Sitemap last modified: 2026-09-13T20:18:50.000Z docs · fts · lemmatization # Lemmatization (engine, not yet exposed). implemented in engine, not yet exposed Lemmatization exists in the OriginChainDB engine, but it is not yet selectable through the HTTP API. There is no language or analyzer query parameter today. The analyzer the FTS API actually applies is Unicode tokenize + lowercase, and nothing else - no stemming and no lemmatization. Concretely: `q=ran` will not match a document that says "running" through today's API. This page describes what lemmatization will do once it is exposed; treat it as roadmap, not as something you can call now. Lemmatization reduces every inflected form to its canonical lemma using a dictionary look-up. Unlike Snowball stemming (a deterministic suffix-strip), lemmatization handles irregular forms — "ran" → "run", "geese" → "goose", "better" → "good". The trade is bigger memory (the dictionary) for higher precision on natural-language fields. ## Stem vs lemma — worked example. inputSnowball stemlemma `running` `run` `run` `ran` `ran` `run` `runs` `run` `run` `better` `bet` `good` `geese` `gees` `goose` `mice` `mice` `mouse` Notice "better" and "geese" — Snowball can't handle either correctly. Lemma does, because the dictionary maps them directly to the canonical form. Both columns are what the engine computes internally; neither is reachable from the API today. ## 9 languages (planned). EnglishSpanishFrenchGermanItalianPortugueseRussianDutchSwedish These are the languages the engine has lemma dictionaries for. Once the analyzer becomes API-selectable, coverage will track Snowball stemming (18 languages) over time. Until then, none of these can be turned on from a query. ## What you can rely on today. Until lemmatization is exposed, the FTS API normalises text in exactly two steps, applied identically at index time and query time: - Unicode tokenize — split into words by Unicode word boundaries (UAX #29). - Lowercase — so `Wireless` matches `wireless`. - The one customer-controlled transform that is wired through is synonyms — install a per-(table, field) map via `POST /fts/:t/:f/synonyms` and members of a class match each other. If you need "ran" to find "run" today, model it as a synonym class, not as lemmatization. ## Where lemma will help (once exposed). - Natural-language fields (article bodies, support tickets, product descriptions) where false positives from stem-collisions hurt precision. - You want "ran" and "running" to match "run" — Snowball can't handle the irregular form, and today's API matches neither. - Field is searchable by humans, not by structured tokens like SKUs or log codes. --- # Graph queries and Cypher on OriginChainDB - BFS, paths, PageRank Canonical source: https://originchaindb.com/docs/graph Sitemap last modified: 2026-09-23T08:33:23.000Z reference · graph # Graph queries Traverse declared relationships between rows to find neighbors, follow paths, and explore connected records. Edges are not stored separately - they're a view over your existing row columns. Writing a row creates / updates / removes the edges automatically. See [Schemas → relations](https://originchaindb.com/docs/schemas/reference#relations) for the declaration. ## 1. Declare a relation. Add a `[[relations]]` block to the schema of the table that holds the foreign-key column. After registering this schema, every row write creates the edges automatically. ``` # manifest.toml - declare the relation on the table holding the FK column. # The reverse edge is written atomically when bidirectional = true. namespace = "social" table = "follows" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "follower_id" ty = "str" required = true [[columns]] name = "followee_id" ty = "str" required = true [[relations]] name = "followee" # the verb you walk in queries from_col = "followee_id" # column on THIS table bidirectional = true [relations.target] namespace = "social" table = "users" pk = "id" ``` Self-relations work - direction tags in the key prevent collisions. See [Schemas → relations](https://originchaindb.com/docs/schemas/reference#relations) for the full reference. ## 2. One-hop neighbors. what this does Return the primary keys of every row directly connected to a given row through a relation. The fastest graph query - just an adjacency list lookup. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.follows/neighbors?rel=followee&pk=f001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` neighbors = db.graph.neighbors( "social.follows", rel="followee", pk="f001", ) for n in neighbors: print(n.pk) ``` #### TypeScript ``` const neighbors = await db.graph.neighbors("social.follows", { rel: "followee", pk: "f001", }); ``` #### Go ``` neighbors, err := db.Graph().Neighbors(ctx, "social.follows", originchain.NeighborsRequest{ Rel: "followee", PK: "f001", }) ``` For inbound edges (who points at this row?), use the parallel `/reverse` endpoint with the same params. ## 3. BFS - multi-hop traversal. what this does Walk every node reachable within `max_depth` hops, breadth-first. Each result carries the hop distance. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.follows/bfs?rel=followee&pk=u001&max_depth=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` hits = db.graph.bfs( "social.follows", rel="followee", pk="u001", max_depth=3, ) for h in hits: print(h.pk, h.depth) ``` #### TypeScript ``` const hits = await db.graph.bfs("social.follows", { rel: "followee", pk: "u001", max_depth: 3, }); ``` #### Go ``` hits, err := db.Graph().BFS(ctx, "social.follows", originchain.BFSRequest{ Rel: "followee", PK: "u001", MaxDepth: 3, }) ``` common mistakes - Forgetting max_depth. On a connected graph, BFS without a depth cap can return millions of nodes. Set `max_depth` to the smallest value that gives you what you need. - Using BFS when you just want reachability. If you only need to know whether two nodes are connected, use `/path` - it short-circuits on first match. ## 4. Shortest path (Dijkstra). what this does Find the lowest-cost path from `src` to `dst`. You provide a map of edge weights; the engine returns the total cost (or `null` if unreachable). #### cURL ``` # weights_json maps "from_pk|to_pk" → cost. Empty {} = every edge weight 1. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/social.follows/dijkstra?rel=followee&src=u001&dst=u042&weights_json=%7B%7D" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` result = db.graph.dijkstra( "social.follows", rel="followee", src="u001", dst="u042", weights={}, # default weight 1 per edge ) print(result.cost) # None if unreachable ``` #### TypeScript ``` const result = await db.graph.dijkstra("social.follows", { rel: "followee", src: "u001", dst: "u042", weights: {}, }); console.log(result.cost); // null if unreachable ``` #### Go ``` result, err := db.Graph().Dijkstra(ctx, "social.follows", originchain.DijkstraRequest{ Rel: "followee", Src: "u001", Dst: "u042", Weights: map[string]float64{}, }) ``` For top-K paths instead of the single shortest, use `/k-shortest` (Yen's algorithm). For unweighted shortest path, use `/path`. ## 5. All 22 endpoints. Every graph endpoint, its wire shape, and what it's good for. Centrality and community algorithms (betweenness, eigenvector, label-propagation, triangles, louvain) require same-schema relations - cross-schema is not yet supported. | operation | signature | use | | --- | --- | --- | | neighbors | GET /graph/:schema/neighbors?rel=&pk= | One-hop forward - downstream of a node. | | reverse | GET /graph/:schema/reverse?rel=&pk= | One-hop inbound - who points TO this node? | | bfs | GET /graph/:schema/bfs?rel=&pk=&max_depth= | Breadth-first frontier up to a depth. | | path | GET /graph/:schema/path?rel=&src=&dst=&max_depth= | Reachability check - short-circuits on first match. | | all_simple_paths | GET /graph/:schema/all_simple_paths?rel=&src=&dst=&max_depth=&max_paths= | Every acyclic (simple) route between two nodes. Default cap 256 paths. | | all_simple_paths_bidir | GET /graph/:schema/all_simple_paths_bidir?rel=&src=&dst=&max_depth=&max_paths= | Same, walking each hop in either direction. Needs bidirectional = true. | | dijkstra | GET /graph/:schema/dijkstra?rel=&src=&dst=&weights_json= | Weighted shortest path. | | k-shortest | GET /graph/:schema/k-shortest?rel=&source=&target=&k= | Top-K loop-free paths in increasing weight. | | pagerank | GET /graph/:schema/pagerank?rel=&nodes= | Influence ranking via power iteration. | | triangles | GET /graph/:schema/triangles?rel= | Per-node triangle count. Clustering signal. | | components | GET /graph/:schema/components?rel= | Connected components via Union-Find. | | louvain | GET /graph/:schema/louvain?rel=&tolerance=&max_levels= | Modularity-greedy community detection. | | betweenness | GET /graph/:schema/betweenness?rel=&max_nodes= | Brandes' betweenness - bridge identification. | | eigenvector_centrality | GET /graph/:schema/eigenvector_centrality?rel=&max_iter=&tol= | Influence by who-you're-connected-to. | | label_propagation | GET /graph/:schema/label_propagation?rel=&max_iter=&seed= | Fast soft community detection. | | random-walk | GET /graph/:schema/random-walk?rel=&start=&steps=&seed=&p=&q= | Uniform or Node2Vec-biased walks. | | node2vec | POST /graph/:schema/node2vec { rel, dim, ..., persist } | Train graph embeddings. | | node2vec/topk | GET /graph/:schema/node2vec/:rel/topk?query=&k=&metric= | Find similar nodes via Node2Vec embeddings. | | graphsage | POST /graph/:schema/graphsage { rel, dim, ..., feature_col, persist } | Attribute-aware embeddings (Hamilton 2017). | | graphsage/topk | GET /graph/:schema/graphsage/:rel/topk?query=&k=&metric= | Find similar nodes via GraphSAGE. | | graphsage/health | GET /graph/:schema/graphsage/:rel/health | Is the persisted embedding set degenerate? Read-only. | | graphsage/rebuild | POST /graph/:schema/graphsage/:rel/rebuild { feature_col, force } | Retrain a stuck index from the rows already stored. | ## Examples. ## One hop. `/neighbors` is the primitive everything else is built on. All four languages wrap it. #### cURL ``` # Who placed this order? curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/neighbors?rel=placed_by&pk=01JTRX9KQ3YH8K2WMX0F5JZAB7" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` const pks = await db.graph.neighbors("shop.orders", { rel: "placed_by", pk: "01JTRX9KQ3YH8K2WMX0F5JZAB7", }); console.log(pks); // string[] - raw primary keys ``` #### Python ``` hits = db.graph.neighbors( "shop.orders", rel="placed_by", pk="01JTRX9KQ3YH8K2WMX0F5JZAB7", ) for n in hits: print(n.pk, n.depth) # depth is always 1 here ``` #### Go ``` hits, err := db.Graph().Neighbors(ctx, "shop.orders", originchain.NeighborsRequest{ Rel: "placed_by", PK: "01JTRX9KQ3YH8K2WMX0F5JZAB7", }) if err != nil { /* handle */ } for _, n := range hits { fmt.Println(n.PK, n.Depth) // Depth is always 1 here } ``` response ``` ["01JTRX1H4Q9P0N2WMX0F5JZ001"] ``` Note what comes back: primary keys, not rows. The graph endpoints return identity, not content. If you want the customer's name you either fetch the row afterwards, or use Cypher / a plan query, which return full rows. ### Going the other way The relation is declared on `shop.orders`, and it stays declared there no matter which direction you walk. `/reverse` is still addressed as `graph/shop.orders/…`, but the `pk` you pass is a customer. #### cURL ``` # Flip it: which orders did this customer place? curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/reverse?rel=placed_by&pk=01JTRX1H4Q9P0N2WMX0F5JZ001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` const pks = await db.graph.reverseNeighbors("shop.orders", { rel: "placed_by", pk: "01JTRX1H4Q9P0N2WMX0F5JZ001", }); console.log(pks.length, "orders for this customer"); ``` #### Python ``` hits = db.graph.reverse_neighbors( "shop.orders", rel="placed_by", pk="01JTRX1H4Q9P0N2WMX0F5JZ001", ) print(len(hits), "orders for this customer") ``` #### Go ``` hits, err := db.Graph().ReverseNeighbors(ctx, "shop.orders", originchain.NeighborsRequest{ Rel: "placed_by", PK: "01JTRX1H4Q9P0N2WMX0F5JZ001", }) if err != nil { /* handle */ } fmt.Println(len(hits), "orders for this customer") ``` empty is not the same as wrong `/reverse` returns `[]` - not an error - in three different situations: the node genuinely has no in-edges, the relation was declared `bidirectional = false`, or you typo'd the `rel` name. Reverse lookups do not validate that the relation exists. Check your spelling before concluding the graph is empty. ## Many hops. `/bfs` expands outward from a node and tags every result with its distance. This is the workhorse for "everything within N hops". #### cURL ``` # Everyone within 3 referral hops of this customer. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.customers/bfs?rel=referrer&pk=01JTRX1H4Q9P0N2WMX0F5JZ001&max_depth=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` const hits = await db.graph.bfs("shop.customers", { rel: "referrer", pk: "01JTRX1H4Q9P0N2WMX0F5JZ001", max_depth: 3, }); for (const h of hits) console.log(h.depth, h.pk); ``` #### Python ``` hits = db.graph.bfs( "shop.customers", rel="referrer", pk="01JTRX1H4Q9P0N2WMX0F5JZ001", max_depth=3, ) for h in hits: print(h.depth, h.pk) ``` #### Go ``` hits, err := db.Graph().BFS(ctx, "shop.customers", originchain.BFSRequest{ Rel: "referrer", PK: "01JTRX1H4Q9P0N2WMX0F5JZ001", MaxDepth: 3, }) if err != nil { /* handle */ } for _, h := range hits { fmt.Println(h.Depth, h.PK) } ``` response ``` [ { "pk": "01JTRX1H4Q9P0N2WMX0F5JZ004", "depth": 1 }, { "pk": "01JTRX1H4Q9P0N2WMX0F5JZ011", "depth": 2 }, { "pk": "01JTRX1H4Q9P0N2WMX0F5JZ027", "depth": 3 } ] ``` ### Reachability and routes `/path` answers one question - can I get there - and answers it cheaply. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.customers/path?rel=referrer&src=01JTRX1H4Q9P0N2WMX0F5JZ001&dst=01JTRX1H4Q9P0N2WMX0F5JZ027&max_depth=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` const res = await db.graph.path("shop.customers", { rel: "referrer", src: "01JTRX1H4Q9P0N2WMX0F5JZ001", dst: "01JTRX1H4Q9P0N2WMX0F5JZ027", max_depth: 3, }); console.log(res.reachable); // boolean - no node list ``` #### Python ``` res = db.graph.path( "shop.customers", rel="referrer", src="01JTRX1H4Q9P0N2WMX0F5JZ001", dst="01JTRX1H4Q9P0N2WMX0F5JZ027", max_depth=3, ) print(res.reachable) # True / False - no node list ``` #### Go ``` res, err := db.Graph().Path(ctx, "shop.customers", originchain.PathRequest{ Rel: "referrer", Src: "01JTRX1H4Q9P0N2WMX0F5JZ001", Dst: "01JTRX1H4Q9P0N2WMX0F5JZ027", MaxDepth: 3, }) if err != nil { /* handle */ } fmt.Println(res.Reachable) // bool - no node list ``` /path does not return a path The response is `{ "reachable": true }` and nothing else. The route is not materialised. If you need the actual nodes, use `/k-shortest` with `k=1`, which does return a node list and a cost. ``` # Want the actual route, with costs? Use k-shortest. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.customers/k-shortest?rel=referrer&source=01JTRX1H4Q9P0N2WMX0F5JZ001&target=01JTRX1H4Q9P0N2WMX0F5JZ027&k=3" \ -H "Authorization: Bearer $OC_TOKEN" ``` response ``` { "paths": [ { "nodes": ["...001", "...004", "...027"], "cost": 2.0 }, { "nodes": ["...001", "...011", "...019", "...027"], "cost": 3.0 } ] } ``` Weighting differs between the two weighted endpoints, and it trips people up. `/dijkstra` takes a `weights_json` map keyed by `"from_pk|to_pk"` that you supply in the request - weights are not stored on edges. `/k-shortest` is usually what you want instead: pass `weight_col` and it reads the weight from a column on the destination row, defaulting to 1.0 per hop. ## Traversal that returns rows: RelationHop. Underneath the REST endpoints, a hop is a query-plan operator called `RelationHop`. It reads the forward edge index by prefix and then point-gets each destination row - so unlike `/neighbors`, it hands back complete rows. You can post a plan directly to `/v1/tenants/:t/query`. The plan is plain JSON, tagged by `op`: ``` { "op": "relation_hop", "schema": "shop.orders", "rel": "placed_by", "from_pk": ["01JTRX9KQ3YH8K2WMX0F5JZAB7"], "target": "shop.customers" } ``` #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/query" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "op": "relation_hop", "schema": "shop.orders", "rel": "placed_by", "from_pk": ["01JTRX9KQ3YH8K2WMX0F5JZAB7"], "target": "shop.customers" }' ``` #### TypeScript ``` const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/query`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ op: "relation_hop", schema: "shop.orders", rel: "placed_by", from_pk: ["01JTRX9KQ3YH8K2WMX0F5JZAB7"], target: "shop.customers", }), }, ); const rows = await res.json(); // full target rows, not just PKs ``` #### Python ``` import os, requests BASE = f"https://{os.environ['OC_HOST']}/v1/tenants/{os.environ['OC_TENANT']}" H = {"Authorization": f"Bearer {os.environ['OC_TOKEN']}"} rows = requests.post(f"{BASE}/query", headers=H, json={ "op": "relation_hop", "schema": "shop.orders", "rel": "placed_by", "from_pk": ["01JTRX9KQ3YH8K2WMX0F5JZAB7"], "target": "shop.customers", }).json() print(rows) # full target rows, not just PKs ``` #### Go ``` plan := map[string]any{ "op": "relation_hop", "schema": "shop.orders", "rel": "placed_by", "from_pk": []string{"01JTRX9KQ3YH8K2WMX0F5JZAB7"}, "target": "shop.customers", } body, _ := json.Marshal(plan) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/query", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var rows []map[string]any // full target rows, not just PKs json.NewDecoder(resp.Body).Decode(&rows) ``` ### Filtering mid-traversal Chaining hops uses a sibling operator, `RelationChain`, and this is where plans earn their keep: each hop can carry its own predicate, applied to that hop's output before the next hop expands from it. That is push-down filtering, not post-filtering - a selective predicate early in the chain cuts the work everything after it does. ``` { "op": "relation_chain", "from_schema": "shop.orders", "from_pk": ["01JTRX9KQ3YH8K2WMX0F5JZAB7"], "hops": [ { "rel": "placed_by", "target": "shop.customers" }, { "rel": "referrer", "target": "shop.customers", "where_predicate": { "op": "eq", "path": "country", "value": "NG" } } ] } ``` Only the final hop's rows come back. And because these are ordinary plan nodes, the usual operators compose around them - filter, project, sort, limit, distinct, aggregate. cycles fan out exponentially `RelationChain` does not de-duplicate rows between hops and carries no built-in depth cap. On a graph with cycles, `A → B → A → …` expands combinatorially with every hop you add. Cap the depth yourself, keep chains short, and prefer `/bfs` - which does track visited nodes - when the shape is "everything within N hops". Most people should reach for [Cypher](https://originchaindb.com/docs/schemas/cypher) rather than hand-writing plans - `MATCH (o:orders {id: "..."})-[:placed_by]->(c) RETURN c.name` compiles to exactly the plan above. Raw plans are there for generated queries and for shapes Cypher does not express. ## Graph algorithms. These are real, callable endpoints - not patterns you assemble yourself. Twenty-two of them, each a single HTTP call against a declared relation. Two worked examples first, then the full catalogue. ### PageRank Ranks influence within a set of nodes you nominate. The `nodes` parameter is required - PageRank scores a subgraph you define, it does not rank an entire table on its own, and an empty list is a `400`. #### cURL ``` # PageRank needs an explicit node universe - it does not rank a whole # table for you. Pass the PKs you want scored. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.customers/pagerank?rel=referrer&nodes=c001,c002,c003,c004,c005&damping=0.85" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` // The TS SDK wraps neighbors / reverse / bfs / path / dijkstra only. // The algorithm endpoints are plain GETs. const qs = new URLSearchParams({ rel: "referrer", nodes: "c001,c002,c003,c004,c005", damping: "0.85", }); const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/graph/shop.customers/pagerank?${qs}`, { headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}` } }, ); const hits: { pk: string; score: number }[] = await res.json(); for (const h of hits) console.log(h.pk, h.score); ``` #### Python ``` scores = db.graph.pagerank( "shop.customers", rel="referrer", nodes=["c001", "c002", "c003", "c004", "c005"], damping=0.85, ) for pk, score in sorted(scores.items(), key=lambda kv: -kv[1]): print(pk, round(score, 4)) ``` #### Go ``` // The Go SDK wraps Neighbors / ReverseNeighbors / BFS / Path / Dijkstra // only. The algorithm endpoints are plain GETs. q := url.Values{} q.Set("rel", "referrer") q.Set("nodes", "c001,c002,c003,c004,c005") q.Set("damping", "0.85") req, _ := http.NewRequestWithContext(ctx, "GET", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/graph/shop.customers/pagerank?"+q.Encode(), nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var hits []struct { PK string `json:"pk"` Score float64 `json:"score"` } json.NewDecoder(resp.Body).Decode(&hits) for _, h := range hits { fmt.Println(h.PK, h.Score) } ``` response ``` [ { "pk": "c001", "score": 0.3120 }, { "pk": "c004", "score": 0.2455 }, { "pk": "c002", "score": 0.1810 } ] ``` ### Louvain communities Partitions the graph into clusters by modularity. Unlike PageRank it takes the whole relation, and community ids are dense integers starting at zero. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.customers/louvain?rel=referrer" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/graph/shop.customers/louvain?rel=referrer`, { headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}` } }, ); const { communities } = await res.json(); for (const c of communities) console.log(c.community, c.pk); ``` #### Python ``` communities = db.graph.louvain("shop.customers", rel="referrer") # {pk -> community_id}, ids are dense integers from 0 for pk, cid in communities.items(): print(cid, pk) ``` #### Go ``` req, _ := http.NewRequestWithContext(ctx, "GET", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/graph/shop.customers/louvain?rel=referrer", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var out struct { Communities []struct { PK string `json:"pk"` Community int `json:"community"` } `json:"communities"` } json.NewDecoder(resp.Body).Decode(&out) for _, c := range out.Communities { fmt.Println(c.Community, c.PK) } ``` response ``` { "communities": [ { "pk": "c001", "community": 0 }, { "pk": "c002", "community": 0 }, { "pk": "c007", "community": 1 } ] } ``` ### The full catalogue All twenty-two live under `/v1/tenants/:t/graph/:schema/` and take `rel=`. Everything is `GET` except the two embedding trainers and the GraphSAGE rebuild. | Endpoint | Shape | What it's for | | --- | --- | --- | | neighbors | GET ?rel=&pk= | One hop forward. Returns a bare array of primary keys. | | reverse | GET ?rel=&pk= | One hop backward. Needs bidirectional = true. | | bfs | GET ?rel=&pk=&max_depth= | Breadth-first frontier with depths. Default depth 3. | | path | GET ?rel=&src=&dst=&max_depth= | Reachability only - returns { reachable } and no route. | | dijkstra | GET ?rel=&src=&dst=&weights_json= | Cheapest weighted route. Weights come from the request, not storage. | | k-shortest | GET ?rel=&source=&target=&k=&weight_col= | Yen's k loop-free routes, with node lists. Max k = 50. | | all_simple_paths | GET ?rel=&src=&dst=&max_depth=&max_paths= | Every acyclic route between two nodes. Default cap 256 paths. | | all_simple_paths_bidir | GET ?rel=&src=&dst=&max_depth=&max_paths= | Same, searching from both ends. Needs bidirectional = true. | | pagerank | GET ?rel=&nodes=&damping=&max_iter=&tol= | Influence ranking by power iteration. nodes= is required. | | triangles | GET ?rel= | Triangle enumeration, each reported once in canonical order. | | components | GET ?rel= | Connected components by union-find. Undirected interpretation. | | betweenness | GET ?rel=&max_nodes= | Brandes' betweenness - finds bridges. Clamped at 100k nodes. | | eigenvector_centrality | GET ?rel=&max_iter=&tol= | Influence weighted by neighbours' influence. | | label_propagation | GET ?rel=&max_iter=&seed= | Fast community detection. Pass seed= or results are not reproducible. | | louvain | GET ?rel=&tolerance=&max_levels= | Modularity-based communities. Up to 500k nodes. | | random-walk | GET ?rel=&start=&steps=&seed=&p=&q= | Biased random walk sampling. Max 1,000 steps. | | node2vec | POST { rel, dim, walks_per_node, …, persist } | Train structural embeddings. persist: true enables topk. | | node2vec/:rel/topk | GET ?query=&k=&metric= | Nearest nodes by persisted Node2Vec embedding. | | graphsage | POST { rel, dim, layers, feature_col, persist } | Attribute-aware embeddings that read a feature column. | | graphsage/:rel/topk | GET ?query=&k=&metric= | Nearest nodes by persisted GraphSAGE embedding. | | graphsage/:rel/health | GET | Health of a persisted GraphSAGE index. Read-only, no row data. | | graphsage/:rel/rebuild | POST { feature_col, force, dry_run } | Retrain a degenerate index in place. 404 if none, 409 if healthy. | not implemented Named here so you don't go looking: - Strongly connected components (Tarjan / Kosaraju) - components is undirected only - Topological sort - Minimum spanning tree - Clustering coefficient - triangles gives you the raw counts to compute it yourself SDK coverage is uneven. TypeScript and Go wrap the five traversal endpoints - neighbors, reverse, bfs, path, dijkstra - and nothing else. Python additionally wraps k-shortest, shortest-path, random-walk, PageRank, Louvain, label propagation, betweenness, and the Node2Vec / GraphSAGE top-k calls. Everything else is a plain GET, as the tabs above show. ## Limits and gotchas. - Graph is a capability you enable, not one you buy - it is included on every paid configuration at no extra charge. Every `/graph/*` route returns `402` with `{"addon":"graph"}` if the capability is not enabled on your instance. Traversal via Cypher goes through a different route and a different check. - Some caps clamp, some reject. `k` above 50 on k-shortest is a `400` - deliberately, so you budget rather than silently getting fewer paths. Betweenness above its node ceiling clamps instead. Know which one you are relying on. - Ceilings worth writing down: variable-length depth 64, k-shortest k 50, random walk 1,000 steps, betweenness 100,000 nodes, Louvain 500,000 nodes, Node2Vec and GraphSAGE 100,000 nodes and 1,024 dimensions, and an 8 MiB HTTP body cap. - Label propagation is non-deterministic unless you pass `seed`. Without one the server seeds from the clock, and the seed is not echoed back - so you cannot reproduce a run after the fact. Always pass your own. - Connected components is undirected. It unions both endpoints of every edge, so it will not give you strongly connected components on a directed graph. - Graph calls are admission-controlled. They are classed as heavy operations and can return `429` with a `Retry-After` under memory pressure, or `413` when a result exceeds the size budget. Both are protective - retry or narrow the query rather than looping hard. ## 6. Examples. ## 6.1 MATCH and RETURN. The simplest query pins one node and projects some properties off it. A label - `:orders` - is matched case-insensitively against registered table names, and `default_schema` is a full `namespace.table` id that applies only to patterns carrying no label. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH (o:orders {id: \"01JTRX9KQ3YH8K2WMX0F5JZAB7\"}) RETURN o.status, o.amount_cents", "default_schema": "shop.orders" }' ``` #### TypeScript ``` // No SDK wraps /cypher yet - all four tabs call the route directly. const BASE = `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}`; const H = { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }; async function cypher(query: string, defaultSchema?: string, params = {}) { const res = await fetch(`${BASE}/cypher`, { method: "POST", headers: H, body: JSON.stringify({ cypher: query, default_schema: defaultSchema, params, }), }); if (!res.ok) throw new Error(await res.text()); return res.json(); } const out = await cypher( `MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"}) RETURN o.status, o.amount_cents`, "shop.orders", ); console.log(out.kind, out.rows); ``` #### Python ``` # No SDK wraps /cypher yet - all four tabs call the route directly. import os, requests BASE = f"https://{os.environ['OC_HOST']}/v1/tenants/{os.environ['OC_TENANT']}" H = {"Authorization": f"Bearer {os.environ['OC_TOKEN']}"} def cypher(query, default_schema=None, params=None): r = requests.post( f"{BASE}/cypher", headers=H, json={ "cypher": query, "default_schema": default_schema, "params": params or {}, }, ) r.raise_for_status() return r.json() out = cypher( 'MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"}) ' 'RETURN o.status, o.amount_cents', default_schema="shop.orders", ) print(out["kind"], out["rows"]) ``` #### Go ``` // No SDK wraps /cypher yet - all four tabs call the route directly. type cypherReq struct { Cypher string `json:"cypher"` DefaultSchema string `json:"default_schema,omitempty"` Params map[string]any `json:"params,omitempty"` } func cypher(ctx context.Context, req cypherReq) (map[string]any, error) { body, _ := json.Marshal(req) hreq, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/cypher", bytes.NewReader(body)) hreq.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) hreq.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(hreq) if err != nil { return nil, err } defer resp.Body.Close() var out map[string]any err = json.NewDecoder(resp.Body).Decode(&out) return out, err } out, err := cypher(ctx, cypherReq{ Cypher: `MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"}) RETURN o.status, o.amount_cents`, DefaultSchema: "shop.orders", }) if err != nil { /* handle */ } fmt.Println(out["kind"], out["rows"]) ``` response ``` { "kind": "select", "rows": [ { "status": "paid", "amount_cents": 12950 } ] } ``` RETURN n gives you an empty object `RETURN o` parses, runs, returns `200`, and hands back `{}` for every row. Whole-node return is not implemented - the projection looks for a column literally named `o` and finds none. Always name the properties you want. The exceptions are variables bound by `WITH … AS n` or `UNWIND … AS n`, which do carry values. Note the column names in that response. `o.status` came back as `status` - the variable prefix is dropped, and the rest of the dotted path becomes the key. Use `AS` when you want to control it. ## 6.2 Filtering. `WHERE` supports `=`, `<>` (and `!=`), `<`, `<=`, `>`, `>=`, `AND`, `OR`, `NOT`, and `IS NULL` / `IS NOT NULL`. Each comparison puts a property on one side and a literal on the other. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH (o:orders) WHERE o.status = \"paid\" AND o.amount_cents > 5000 RETURN o.id, o.amount_cents", "default_schema": "shop.orders" }' ``` #### TypeScript ``` const out = await cypher( `MATCH (o:orders) WHERE o.status = "paid" AND o.amount_cents > 5000 RETURN o.id, o.amount_cents`, "shop.orders", ); ``` #### Python ``` out = cypher( 'MATCH (o:orders) ' 'WHERE o.status = "paid" AND o.amount_cents > 5000 ' 'RETURN o.id, o.amount_cents', default_schema="shop.orders", ) ``` #### Go ``` out, err := cypher(ctx, cypherReq{ Cypher: `MATCH (o:orders) WHERE o.status = "paid" AND o.amount_cents > 5000 RETURN o.id, o.amount_cents`, DefaultSchema: "shop.orders", }) ``` Three limits to internalise now, because the error messages for them are generic. There is no `IN`, no `STARTS WITH` / `CONTAINS` / `ENDS WITH`, and no regex - those all surface as `cypher parse: trailing tokens after query`, which reads like a syntax slip rather than a missing feature. And you cannot compare two nodes to each other: `WHERE o.customer = c.id` is refused. ## 6.3 Walking one hop. `-[:placed_by]->` names the relation declared on `shop.orders`. The engine reads the pre-built edge index rather than scanning the target table, so a hop costs a prefix lookup plus one point-get per neighbour. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH (o:orders {id: \"01JTRX9KQ3YH8K2WMX0F5JZAB7\"})-[:placed_by]->(c) RETURN c.name, c.country", "default_schema": "shop.orders" }' ``` #### TypeScript ``` const out = await cypher( `MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"})-[:placed_by]->(c) RETURN c.name, c.country`, "shop.orders", ); ``` #### Python ``` out = cypher( 'MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"})-[:placed_by]->(c) ' 'RETURN c.name, c.country', default_schema="shop.orders", ) ``` #### Go ``` out, err := cypher(ctx, cypherReq{ Cypher: `MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"})-[:placed_by]->(c) RETURN c.name, c.country`, DefaultSchema: "shop.orders", }) ``` response ``` { "kind": "select", "rows": [ { "name": "Ada Okafor", "country": "NG" } ] } ``` ### Backwards, and both ways `<-[:rel]-` walks the edge in reverse, and `-[:rel]-` walks it in both. Reverse traversal only works when the relation declares `bidirectional = true` - which is the default, but a relation explicitly set to `false` silently returns nothing rather than erroring. ``` # Which orders did this customer place? Walk the edge backwards. # Only legal because placed_by declares bidirectional = true. MATCH (c:customers {id: "01JTRX1H4Q9P0N2WMX0F5JZ001"})<-[:placed_by]-(o) RETURN o.id, o.amount_cents ``` every traversal needs an anchor The first node of any pattern that contains a hop must pin its primary key in the property map. `MATCH (o:orders)-[:placed_by]->(c)` is a hard error, not a slow full-graph query - the message asks you to add `{id: ...}`. A bare `MATCH (o:orders)` with no hop is fine and scans. ## 6.4 Walking several hops. Chain arrows to cross more than one relation in a single pattern. Each hop names its own relation, and only the last hop's nodes come back as rows. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH (o:orders {id: \"01JTRX9KQ3YH8K2WMX0F5JZAB7\"})-[:placed_by]->(c)-[:referrer]->(r) RETURN r.name, r.country", "default_schema": "shop.orders" }' ``` #### TypeScript ``` // order -> customer -> the customer who referred them const out = await cypher( `MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"})-[:placed_by]->(c)-[:referrer]->(r) RETURN r.name, r.country`, "shop.orders", ); ``` #### Python ``` # order -> customer -> the customer who referred them out = cypher( 'MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"})' '-[:placed_by]->(c)-[:referrer]->(r) ' 'RETURN r.name, r.country', default_schema="shop.orders", ) ``` #### Go ``` // order -> customer -> the customer who referred them out, err := cypher(ctx, cypherReq{ Cypher: `MATCH (o:orders {id: "01JTRX9KQ3YH8K2WMX0F5JZAB7"}) -[:placed_by]->(c)-[:referrer]->(r) RETURN r.name, r.country`, DefaultSchema: "shop.orders", }) ``` ### Variable-length paths When you don't know the depth in advance, `-[:rel*1..3]->` walks between one and three hops and returns everything it reaches. Bind the path with `p =` to use `length(p)`, `nodes(p)` and `relationships(p)`. This only works on a self-recursive relation - one whose target table declares a relation of the same name, like `referrer` on `shop.customers`. You cannot walk a variable number of hops across `placed_by`, because orders do not point at orders. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH p = (c:customers {id: \"01JTRX1H4Q9P0N2WMX0F5JZ001\"})-[:referrer*1..3]->(up) RETURN up.name, length(p)", "default_schema": "shop.customers" }' ``` #### TypeScript ``` // Everyone up to 3 referral hops above this customer. const out = await cypher( `MATCH p = (c:customers {id: "01JTRX1H4Q9P0N2WMX0F5JZ001"})-[:referrer*1..3]->(up) RETURN up.name, length(p)`, "shop.customers", ); ``` #### Python ``` # Everyone up to 3 referral hops above this customer. out = cypher( 'MATCH p = (c:customers {id: "01JTRX1H4Q9P0N2WMX0F5JZ001"})' '-[:referrer*1..3]->(up) ' 'RETURN up.name, length(p)', default_schema="shop.customers", ) ``` #### Go ``` // Everyone up to 3 referral hops above this customer. out, err := cypher(ctx, cypherReq{ Cypher: `MATCH p = (c:customers {id: "01JTRX1H4Q9P0N2WMX0F5JZ001"}) -[:referrer*1..3]->(up) RETURN up.name, length(p)`, DefaultSchema: "shop.customers", }) ``` response ``` { "kind": "select", "rows": [ { "name": "Ravi Menon", "length(p)": 1 }, { "name": "Sofia Duarte","length(p)": 2 } ] } ``` A bare `*` means `*1..64`; 64 is the hard depth ceiling. Two shapes are refused inside a longer chain: a variable-length hop cannot be one link of a multi-hop pattern, and a chain may contain at most three undirected hops (each one doubles the work). `shortestPath(…)` is available over a variable-length pattern when you only want the shortest route between two pinned nodes. ## 6.5 Ordering and limiting. `ORDER BY`, `SKIP` and `LIMIT` attach to `RETURN` and nowhere else - a `WITH … ORDER BY` in the middle of a query does not parse. Sort keys must be a property access or a bare variable; `ASC` is the default. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH (o:orders) WHERE o.status = \"paid\" RETURN o.id AS order_id, o.amount_cents AS cents ORDER BY o.amount_cents DESC SKIP 0 LIMIT 10", "default_schema": "shop.orders" }' ``` #### TypeScript ``` const out = await cypher( `MATCH (o:orders) WHERE o.status = "paid" RETURN o.id AS order_id, o.amount_cents AS cents ORDER BY o.amount_cents DESC SKIP 0 LIMIT 10`, "shop.orders", ); ``` #### Python ``` out = cypher( 'MATCH (o:orders) WHERE o.status = "paid" ' 'RETURN o.id AS order_id, o.amount_cents AS cents ' 'ORDER BY o.amount_cents DESC ' 'SKIP 0 LIMIT 10', default_schema="shop.orders", ) ``` #### Go ``` out, err := cypher(ctx, cypherReq{ Cypher: `MATCH (o:orders) WHERE o.status = "paid" RETURN o.id AS order_id, o.amount_cents AS cents ORDER BY o.amount_cents DESC SKIP 0 LIMIT 10`, DefaultSchema: "shop.orders", }) ``` `SKIP` and `LIMIT` take integer literals. A parameter there does not parse, so build the number into the query string when you paginate. ## 6.6 Parameters. `$name` placeholders are substituted from the request's `params` map before the query is planned. Values must be scalars - string, number, boolean or null; arrays and objects are rejected. Referencing a parameter you didn't supply is a `400`, not a silent null. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH (o:orders) WHERE o.status = $want RETURN o.id", "default_schema": "shop.orders", "params": { "want": "paid" } }' ``` #### TypeScript ``` const out = await cypher( `MATCH (o:orders) WHERE o.status = $want RETURN o.id`, "shop.orders", { want: "paid" }, ); ``` #### Python ``` out = cypher( "MATCH (o:orders) WHERE o.status = $want RETURN o.id", default_schema="shop.orders", params={"want": "paid"}, ) ``` #### Go ``` out, err := cypher(ctx, cypherReq{ Cypher: `MATCH (o:orders) WHERE o.status = $want RETURN o.id`, DefaultSchema: "shop.orders", Params: map[string]any{"want": "paid"}, }) ``` Parameters work in `WHERE`, `RETURN` and `UNWIND`. They do not work inside a pattern's property map, so the very place you most want one - `(o {id: $id})` - is refused, and the anchor id has to be interpolated into the query text. Escape it yourself. ## 6.7 Writing through Cypher. All four write verbs are implemented and go through the same route. Each returns its own response shape rather than rows. #### CREATE ``` # CREATE - insert one node. Returns {"kind":"insert","rows_inserted":1} CREATE (o:orders { id: "01JTRXNEW00000000000000001", customer: "01JTRX1H4Q9P0N2WMX0F5JZ001", amount_cents: 4200, status: "pending", notes: "gift wrap", placed_ms: 1714480000000 }) ``` #### MERGE ``` # MERGE - insert only if absent. Returns {"kind":"merge","rows_inserted":0|1} MERGE (c:customers { id: "01JTRX1H4Q9P0N2WMX0F5JZ009", name: "Nia Blake", country: "GB" }) ``` #### SET ``` # SET - update properties on an anchored node. # Returns {"kind":"update","rows_affected":1} MATCH (o:orders {id: "01JTRXNEW00000000000000001"}) SET o.status = "paid" ``` #### DELETE ``` # DELETE - retire an anchored node and its derived index/edge state. # Returns {"kind":"delete","rows_deleted":1} MATCH (o:orders {id: "01JTRXNEW00000000000000001"}) DELETE o ``` Sent over the wire, an update looks like any other Cypher call: #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/cypher" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "cypher": "MATCH (o:orders {id: \"01JTRXNEW00000000000000001\"}) SET o.status = \"paid\"", "default_schema": "shop.orders" }' ``` #### TypeScript ``` const out = await cypher( `MATCH (o:orders {id: "01JTRXNEW00000000000000001"}) SET o.status = "paid"`, "shop.orders", ); ``` #### Python ``` out = cypher( 'MATCH (o:orders {id: "01JTRXNEW00000000000000001"}) SET o.status = "paid"', default_schema="shop.orders", ) ``` #### Go ``` out, err := cypher(ctx, cypherReq{ Cypher: `MATCH (o:orders {id: "01JTRXNEW00000000000000001"}) SET o.status = "paid"`, DefaultSchema: "shop.orders", }) ``` #### Go ``` map[kind:update rows_affected:1] ``` #### Python ``` {"kind": "update", "rows_affected": 1} ``` #### TypeScript ``` { kind: "update", rows_affected: 1 } ``` `SET` and `DELETE` both require a preceding `MATCH` that pins a primary key. An unanchored mutation would be a whole-table rewrite, so it is refused outright. A trailing `RETURN` after a write parses but is ignored - the write's own counter is what you get back. A constraint violation returns `409` with `{"error":"constraint_violation","detail":"…"}`. ## Result shapes. Every response carries a `kind` discriminator; branch on it. There is no `columns` array and no `data` envelope - reads hand back plain JSON objects keyed by the projection's output names. | kind | Body | Emitted by | | --- | --- | --- | | select | rows: [ … ] | Any query ending in RETURN | | insert | rows_inserted: n | CREATE, and FOREACH bodies | | merge | rows_inserted: n | MERGE - counts only rows that were absent | | update | rows_affected: n | SET | | delete | rows_deleted: n | DELETE | Every response also carries an `X-OC-Query-Id` header - log it, it is what support will ask for. ## Every clause that works. | Clause | Notes | | --- | --- | | MATCH | Node patterns, relationship patterns, chained hops, and comma-separated patterns. | | OPTIONAL MATCH | Left-outer semantics. Cannot be the first clause of a query - put a MATCH before it. | | WHERE | After MATCH, OPTIONAL MATCH or WITH. Property-vs-literal comparisons only. | | RETURN / RETURN DISTINCT | Projections and aliases. See the RETURN n warning above. | | ORDER BY / SKIP / LIMIT | Only inside RETURN. SKIP and LIMIT take non-negative integer literals, not parameters. | | WITH | Projection between stages. WITH DISTINCT and aggregates inside WITH are both refused. | | UNWIND | Expands a literal list, or a list column on a matched row, into rows. | | CREATE | Inserts a node. Can also set a relationship column when it follows a MATCH. | | MERGE | Insert-if-absent on a node pattern. Existing rows are left untouched. | | SET | var.prop = literal, on a node anchored by primary key. | | DELETE | One variable, anchored by primary key. Removes derived index and edge state too. | | FOREACH | Write-only bodies: CREATE, SET, DELETE, or a nested FOREACH. | | CALL { … } YIELD | Subquery form only, read-only, no nesting. | | shortestPath(…) | Over a variable-length pattern between two primary-key-anchored nodes. | Aggregates - `count(*)`, `count(x)`, `sum`, `avg`, `min`, `max` - work, but only on their own: there is no `GROUP BY`, so an aggregate cannot share a `RETURN` with a plain column. The scalar functions `upper`, `lower`, `length`, `coalesce` and `abs` are also available. ## Every clause that does not. This is the list people arrive expecting. Most of these fail with a generic parse error rather than a helpful one, so check here before assuming you have a syntax bug. | Not supported | What to do instead | | --- | --- | | UNION | Refused with a pointer to SQL UNION on /sql, or merge client-side. | | REMOVE | Refused. Use SQL UPDATE ... SET col = NULL instead. | | DETACH DELETE | Parses, then refused. Plain DELETE already clears derived edge state. | | IN | No such operator. Write it as OR-ed equalities. | | STARTS WITH / CONTAINS / ENDS WITH | No string predicates at all. Use full-text search or SQL LIKE. | | =~ (regex) | The ~ character does not even lex - you get a lex error, not a parse error. | | CASE | No conditional expressions. Do it in SQL or in your application. | | EXISTS(…) | Not a recognised function. | | collect() | Not implemented. count / sum / avg / min / max are. | | count(DISTINCT x) | DISTINCT inside an aggregate is rejected. RETURN DISTINCT works. | | RETURN * | Not an expression. List the properties you want. | | Arithmetic in WHERE or RETURN | o.amount_cents * 2 is refused; the error points you at /sql. | | Aggregates mixed with plain columns | There is no GROUP BY in this dialect. Aggregate alone, or use SQL. | | Relationship variables -[r:TYPE]-> | The parser demands a colon straight after the bracket. Edges cannot be bound. | | Untyped edges: --> or -[]-> | Every hop must name a declared relation. | | Edge property maps / type alternation | -[:R {since: 2020}]-> and -[:A\|:B]-> both fail to parse. | | Multi-label nodes (n:A:B) | One label per node pattern. | | Parameters inside pattern maps | (o {id: $x}) is refused - a property map takes literals. $params work in WHERE, RETURN and UNWIND. | | allShortestPaths(…) | Only the singular shortestPath() exists. | | CALL db.labels() and other procedures | Only the CALL { subquery } form is implemented. | | Cross-variable comparison a.x = b.y | WHERE compares a property against a literal, not against another node. | ## Limits and gotchas. - Composite primary keys cannot be used. Every anchored pattern resolves through a single primary-key column. A table with a two-column key is not reachable from Cypher. - Labels resolve to table names, first match wins. If two namespaces both hold a table called `orders`, a bare `:orders` label is ambiguous, and `default_schema` does not disambiguate it - the resolver never consults it once a label is present. - Two MATCH clauses need a WITH between them. Back-to-back MATCHes are refused; project through `WITH` to chain stages. - Comma-separated patterns are a separate, stricter mode. `MATCH (a)-[:r]->(b), (b)-[:r]->(c) RETURN a, b, c` runs a dedicated pattern-matching join, and it accepts only that exact shape: no other clauses, no variable-length, no undirected hops, every node needs a variable, RETURN takes bare node variables only, and WHERE may contain nothing but AND-ed `a <> b` distinctness checks. - A cyclic graph can fan out. Multi-hop chains do not de-duplicate visited rows, so a cycle expands combinatorially with each hop. Keep chains short and prefer a bounded variable-length range. - Large results are refused, not truncated. Exceeding the result-row or result-byte budget returns `413` with the observed size and the cap. Add a `LIMIT`. - Under load you may get `429`. Cypher counts as a heavy operation and is admission-controlled. Honour `Retry-After`. - The request body cap is 8 MiB, which matters if you generate long `UNWIND` literal lists. - Cypher calls are authorised by what they actually do - a pure read is checked like any other SELECT, and each write verb requires the matching privilege on its target schema. A read-only token can run a read-only Cypher query. --- # Graph embeddings on OriginChainDB - Node2Vec & GraphSAGE Canonical source: https://originchaindb.com/docs/graph/embeddings Sitemap last modified: 2026-09-13T20:18:50.000Z docs · graph · embeddings # Graph embeddings. Two embedding families ship in-engine: Node2Vec (biased random-walk) and GraphSAGE (attribute-aware, neighbour-aggregating, three aggregators). Both persist to disk and install into the [vector index](https://originchaindb.com/docs/vector) so you can run nearest-neighbour search over learned node representations. Both train against a relation you have already declared on the schema, so declare the relation first. ## Node2Vec. Skip-gram over biased random walks. Two knobs decide what the walks explore: `p` (return likelihood) and `q` (in-out bias). Low `q` favors depth-first exploration (structural roles); high `q` favors breadth-first (homophily). - `p`Return likelihood. Higher = less revisiting. - `q`In-out bias. < 1 = depth-first; > 1 = breadth-first. - `walk_length`Steps per walk. 20–80 is the usual band. - `num_walks`Walks started per node. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/graph/users/follows/embeddings/node2vec/train" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "dim": 128, "walk_length": 40, "num_walks": 10, "p": 1.0, "q": 0.5, "window": 10 }' ``` ## GraphSAGE. Attribute-aware. Each layer samples a neighbourhood, transforms each neighbour's representation, then aggregates. Three aggregators ship — pick the one that fits your graph shape. aggregatorwhen to pickcost `Mean` Default. Smooths neighbour attributes by averaging. Fastest aggregator; reasonable accuracy on most homophilous graphs. Cheapest. `MaxPool` Element-wise max of a per-neighbour transform. Best representational capacity per parameter; pick when neighbour outliers carry signal. ≈ 2× Mean per training step. `LSTM` Sequence-aware aggregation with deterministic per-(node, layer) shuffle so retrains converge to the same vector. Pick when neighbour ordering matters or the graph carries time-like structure. ≈ 5–8× Mean per training step. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/graph/users/follows/embeddings/graphsage/train" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "dim": 128, "aggregator": "max_pool", // "mean" | "max_pool" | "lstm" "layers": 2, "sample_per_layer": [10, 5] }' ``` ## Persistence + top-k. Trained embeddings persist to disk. Install them into the vector index via the same install-centroids HTTP admin used for IVF, then the engine treats them as a regular vector collection — `topk` gives you nearest-neighbour search over node embeddings, filterable by metadata. See [docs/vector](https://originchaindb.com/docs/vector) for the topk surface. --- # Graph examples by industry - fraud rings - OriginChainDB Canonical source: https://originchaindb.com/docs/graph/industries Sitemap last modified: 2026-09-13T20:18:50.000Z graph · by industry # Graph models by industry Four questions that are awkward in rows and natural as a traversal: a fraud ring, a two-hop product recommendation, a supply-chain blast radius and an org chart. Each starts with the relation, because declaring which column is an edge is the whole setup. ## Financial services: finding a fraud ring Individually ordinary accounts, connected in a way that is not. The signal is the shape, not any single row. ### The relation ``` fin.transfers (id PK, from_account, to_account, amount_minor, at) relation: "to_account" -> fin.accounts ``` ### Who is within three hops of a flagged account A ring shows up as a small set reachable in a few hops and reaching back. One traversal replaces a recursive query you would otherwise write by hand. ``` GET /v1/tenants/:t/graph/fin.transfers/bfs ?rel=to_account&from=acct-flagged&depth=3 ``` keep it explainable A traversal returns the path, not just the endpoint, so an investigator can see WHY an account was surfaced. That matters more than the score when a decision has to be justified. ## Retail: people who bought this also bought The classic two-hop recommendation, computed on live orders rather than a nightly export. ### The relation ``` shop.lines (id PK, order_id, product_id, qty) relation: "product_id" -> shop.products ``` ### Two hops from a product Product to the orders containing it, then out to the other products in those orders. Depth two, one call. ``` GET /v1/tenants/:t/graph/shop.lines/bfs ?rel=product_id&from=sku-8842&depth=2 ``` ### Then rank them by similarity Traversal gives candidates; vector search orders them by how alike they actually are. The two surfaces read the same rows — see [the vector quickstart](https://originchaindb.com/docs/vector/quickstart). ``` POST /v1/tenants/:t/vector/shop.products/topk ``` ## Manufacturing: a supply-chain dependency One supplier stops shipping. The question is what stops with it, several levels down. ### The relation ``` mfg.bom (id PK, part_id, depends_on_part, qty) relation: "depends_on_part" -> mfg.parts ``` ### Everything downstream of a part Blast radius is a traversal, and the depth is the answer to how far the problem travels. ``` GET /v1/tenants/:t/graph/mfg.bom/bfs ?rel=depends_on_part&from=part-4471&depth=5 ``` weights make it a route Give the edge a cost - lead time, price, distance - and Dijkstra answers the cheapest path rather than the shortest hop count. ## Enterprise: an org chart Reporting lines are a graph that everyone models as rows and then queries badly. ### The relation ``` hr.people (id PK, name, manager_id, dept) relation: "manager_id" -> hr.people ``` ### Everyone under a manager A self-referencing relation traverses the same way as any other. ``` GET /v1/tenants/:t/graph/hr.people/bfs ?rel=manager_id&from=p-17&depth=6 ``` ## What these have in common - The relation is the model. Declaring which column is an edge is the whole setup. - No second database. Edges are columns on rows you already write, so there is nothing to sync and nothing to fall behind. - Paths, not just answers. A traversal returns how it got there, which is what makes the result defensible. [Start from zero Declare a relation and traverse it.](https://originchaindb.com/docs/graph/quickstart)[All algorithms The full set, with what each one costs.](https://originchaindb.com/docs/graph) --- # Graph quickstart - neighbors, BFS, paths - OriginChainDB Canonical source: https://originchaindb.com/docs/graph/quickstart Sitemap last modified: 2026-09-23T08:33:23.000Z graph · quickstart # Run your first graph traversal Declare a relationship between tables and traverse the connected rows. ## Before you start An instance, an API key, and a table where one column already holds another table’s key. That column is the edge; nothing is copied into a separate graph store. ``` export OC_URL='https://' export OC_TENANT='' export OC_TOKEN='' ``` ## The four steps 1. 1. Have the rows Any table works. A follow graph is two identifier columns on one table. ``` social.follows (id PK, follower, followee) // follower and followee both name a user id ``` 2. 2. Declare the relation The relation tells the engine which column is an edge. Declared once in the schema — see [graph schema](https://originchaindb.com/docs/schemas/graph) for the full form. ``` [[relations]] name = "followee" column = "followee" target = "social.users" ``` 3. 3. Ask for neighbors One hop from a starting node, along the relation you named. ``` curl "$OC_URL/v1/tenants/$OC_TENANT/graph/social.follows/neighbors?rel=followee&from=u-1" \ -H "Authorization: Bearer $OC_TOKEN" ``` 4. 4. Go further, or find a path Breadth-first search walks several hops; Dijkstra finds the cheapest path when the edges carry a weight. ``` # everyone within three hops curl "$OC_URL/v1/tenants/$OC_TENANT/graph/social.follows/bfs?rel=followee&from=u-1&depth=3" \ -H "Authorization: Bearer $OC_TOKEN" # cheapest route between two nodes curl "$OC_URL/v1/tenants/$OC_TENANT/graph/social.follows/dijkstra?rel=followee&from=u-1&to=u-9" \ -H "Authorization: Bearer $OC_TOKEN" ``` why this is not a second database The edges are columns on rows you already write. There is no separate graph store to load, sync or keep consistent, and a row inserted a second ago is traversable now. ## What to read next [Graph reference Neighbors, BFS, Dijkstra and the full algorithm set.](https://originchaindb.com/docs/graph)[Cypher Pattern-matching syntax over the same relations.](https://originchaindb.com/docs/schemas/cypher)[By industry Fraud rings, recommendations, supply chains, org charts.](https://originchaindb.com/docs/graph/industries) --- # Insert data into OriginChainDB - the complete guide Canonical source: https://originchaindb.com/docs/insert Sitemap last modified: 2026-09-23T08:33:23.000Z how-to · insert # Insert data Write rows, vectors, search documents, and relationships through the HTTP API or an OriginChainDB SDK. New to OriginChainDB? Read [Quickstart](https://originchaindb.com/docs/quickstart) first - it shows you how to create an instance and get the URL + token used below. ## 0. Set up your client. Each language needs three things to talk to OriginChainDB: your endpoint URL, a bearer token, and your tenant ID. Set them up once, then every example below assumes they are in scope. #### cURL (env vars) ``` # Save your endpoint + token + tenant ID once. # You get all three from the dashboard after creating an instance. export ORIGINCHAIN_URL="https://acme.db.originchain.ai" export OC_TOKEN="oc_live_xxxxxxxxxxxxxxxx" export T="acme" # your tenant ID (the part before .db...) ``` #### Python ``` # pip install originchain from originchain import OriginChain db = OriginChain( base_url="https://acme.db.originchain.ai", bearer="oc_live_xxxxxxxxxxxxxxxx", tenant="acme", ) # All examples below assume `db` is in scope. ``` #### TypeScript ``` // npm install @originchain/sdk import { OriginChainClient } from "@originchain/sdk"; const db = new OriginChainClient({ baseUrl: "https://acme.db.originchain.ai", bearer: "oc_live_xxxxxxxxxxxxxxxx", }); // All examples below assume `db` is in scope. // For the few endpoints the SDK does not wrap yet // (row writes, batch writes), we also reuse: const BASE_URL = "https://acme.db.originchain.ai"; const TENANT = "acme"; const OC_TOKEN = "oc_live_xxxxxxxxxxxxxxxx"; ``` #### Go ``` // go get github.com/originchain-ai/originchain-go package main import ( "bytes" "context" "encoding/json" "net/http" "github.com/originchain-ai/originchain-go" ) const ( BASE_URL = "https://acme.db.originchain.ai" TENANT = "acme" OC_TOKEN = "oc_live_xxxxxxxxxxxxxxxx" ) var ctx = context.Background() var db = originchain.NewClient(originchain.Config{ BaseURL: BASE_URL, Bearer: OC_TOKEN, }) // All examples below assume `db` and `ctx` are in scope. ``` where do these come from? - Endpoint URL: Dashboard → your instance → "Connect". Looks like `https://acme.db.originchain.ai`. - Bearer token: Dashboard → your instance → "API tokens" → "Create token". Starts with `oc_live_`. Store it in a secret manager - it grants full access to your instance. - Tenant ID: The first part of the endpoint hostname. For `acme.db...` the tenant is `acme`. SDK status The Python SDK has helpers for every endpoint on this page. The TypeScript and Go SDKs cover vector, full-text, SQL, graph, and ask - but they don't have row-write helpers yet (shipping in the next release). For row writes in TypeScript and Go we show `fetch` / `net/http` calls; they hit the exact same endpoint the SDK will use. ## 1. Insert one row. what this does Save one record to a table - like one row in a spreadsheet. The record is a JSON object whose keys match the column names you declared on the schema. when to use it - You are saving one new record (a user signing up, a single order). - You are updating an existing record - same call. Sending the same `id` again replaces the row. - If you have many rows to save, jump to [Insert many rows](https://originchaindb.com/docs/insert#bulk) - it is much faster. the code #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/shop.customers" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "c_1", "email": "ada@example.com", "region": "IN" }' ``` #### Python ``` db.rows.put("shop.customers", { "id": "c_1", "email": "ada@example.com", "region": "IN", }) ``` #### TypeScript ``` // The TypeScript SDK does not wrap row writes yet // (shipping in the next release). Use `fetch` for now - // it is exactly the same HTTP call the SDK will make. await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/shop.customers`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "c_1", email: "ada@example.com", region: "IN", }), }); ``` #### Go ``` // The Go SDK does not wrap row writes yet // (shipping in the next release). Use net/http for now. body, _ := json.Marshal(map[string]any{ "id": "c_1", "email": "ada@example.com", "region": "IN", }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/shop.customers", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() ``` what each field means | Field | Type | Required | What it is | | --- | --- | --- | --- | | URL `:t` | string | yes | Your tenant ID. The first part of your endpoint hostname. | | URL `:schema` | string | yes | The schema name (here, `shop.customers`). Must already exist - see [Define a schema](https://originchaindb.com/docs/schemas). | | id | string | yes | The primary key declared in your schema. Re-sending the same `id` replaces the row. | | email, region, ... | any | depends | Any other column you declared on the schema. The JSON key matches the column name; the JSON type must match the declared type. | | Authorization | header | yes | `Bearer `. Missing or wrong → `401 unauthorized`. | what you get back ``` { "ok": true, "lsn": { "segment": 4, "offset": 8421007 } } ``` `ok: true` means the row is saved and durable. `lsn` is the commit position where your row landed - useful if you need to wait for a replica to catch up. You can ignore it most of the time. common mistakes - Schema doesn't exist yet. You will see `404 schema_not_found`. Create the schema first - see [Define a schema](https://originchaindb.com/docs/schemas). - Wrong type for a column. If you declared `price` as a number but sent a string, you will see `400 type_mismatch`. The error message names the offending field. - Forgot the Content-Type header. Without `Content-Type: application/json` the server can't parse the body and returns `400 invalid_body`. - Did not mean to overwrite. By default a duplicate `id` overwrites the existing row. If you want the write to fail when the row already exists, add the query string `?expect=insert` - it returns `409 conflict` on duplicate. try it yourself · 30 seconds After [setting up your instance](https://originchaindb.com/docs/quickstart) and creating the `shop.customers` schema, paste this into your terminal: ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/shop.customers" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"id":"test_1","email":"you@example.com","region":"US"}' ``` You should see `{"ok":true,...}`. Run it again - same result. That confirms inserts are idempotent (re-running doesn't break anything). ## 2. Insert many rows. what this does Save many rows in one HTTP call. Far faster than calling "insert one row" in a loop, because each call has fixed network overhead. when to use it - Importing a CSV / JSON file - hundreds, thousands, or millions of rows. - Backfilling a new table from an old one. - Ingesting a stream of events in micro-batches. There are two transport shapes. A JSON array body is simple - send up to ~8 MiB of rows per call. For larger imports, NDJSON (one JSON object per line) lifts that cap and streams. the code #### cURL ``` # 1) JSON array body - send up to ~8 MiB worth of rows in one call. curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/shop.customers/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "c_1", "email": "ada@example.com", "region": "IN" }, { "id": "c_2", "email": "hopper@example.com", "region": "US" }, { "id": "c_3", "email": "lovelace@example.com", "region": "GB" } ]' # 2) NDJSON stream - for millions of rows. No 8 MiB cap. # One JSON object per line. The `chunk` query param controls # how many rows go into each atomic write. curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/shop.customers/_batch?chunk=1000" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/x-ndjson" \ --data-binary @customers.ndjson # customers.ndjson - one row per line: # {"id":"c_1","email":"ada@example.com","region":"IN"} # {"id":"c_2","email":"hopper@example.com","region":"US"} # ... ``` #### Python ``` # put_batch handles the chunking, retries, and idempotency keys. # It accepts any iterable - lists, generators, file-line iterators. def stream_customers(): for i in range(50_000): yield { "id": f"c_{i}", "email": f"user{i}@example.com", "region": "IN", } inserted = db.rows.put_batch( "shop.customers", stream_customers(), chunk=1000, # rows per atomic write idempotency_key="bulk-import-2026-06-10", ) print(f"{inserted} rows accepted") ``` #### TypeScript ``` // Same caveat as a single row insert - use `fetch`. // We pass a JSON array body. await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/shop.customers/_batch`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify([ { id: "c_1", email: "ada@example.com", region: "IN" }, { id: "c_2", email: "hopper@example.com", region: "US" }, { id: "c_3", email: "lovelace@example.com", region: "GB" }, ]), }); ``` #### Go ``` rows := []map[string]any{ {"id": "c_1", "email": "ada@example.com", "region": "IN"}, {"id": "c_2", "email": "hopper@example.com", "region": "US"}, {"id": "c_3", "email": "lovelace@example.com", "region": "GB"}, } body, _ := json.Marshal(rows) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/shop.customers/_batch", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() ``` what each field means | Field | Where | Default | What it controls | | --- | --- | --- | --- | | Body | JSON | - | An array of row objects. Up to ~8 MiB of total request body, roughly 50 000 small rows. | | Body | NDJSON | - | One JSON object per line. No size cap on the body. Use `Content-Type: application/x-ndjson`. | | chunk | query | 1000 | How many rows go into one atomic write. Bigger chunks = fewer durable commits, more throughput. Smaller chunks = finer-grained retry. | | expect | query | - | `?expect=insert` fails the batch if any row already exists. Useful for initial imports where duplicates are bugs. | | Idempotency-Key | header | auto | If you retry the same call with the same key, the server returns the original result without re-inserting. The SDKs set this automatically. | common mistakes - Sending NDJSON with the wrong Content-Type. Use `application/x-ndjson`, not `application/json`. Otherwise the server tries to parse the whole body as one JSON object and fails. - Chunk too big. A single chunk that doesn't fit in memory will OOM the engine's batch buffer. If you're streaming millions of rows, keep `chunk` at the default 1000. - No retry strategy. Network blips happen. The SDKs handle this for you; if you call the endpoint directly, retry on `503` and `504` with exponential backoff. The auto-set `Idempotency-Key` makes the retry safe. try it yourself · 1 minute Generate 1000 fake rows and import them in one call. Save this as `demo.py` and run `python demo.py`: ``` from originchain import OriginChain db = OriginChain.from_env() rows = ({"id": f"c_{i}", "email": f"u{i}@ex.com", "region": "IN"} for i in range(1000)) print(db.rows.put_batch("shop.customers", rows, chunk=500), "rows inserted") ``` ## 3. Insert a vector. what this does Save a list of numbers (an embedding) under an ID so you can later find similar embeddings. An embedding is the output of a model that turned some text or an image into numbers - typically 384, 768, 1024, or 1536 of them. when to use it - You are building semantic search ("find products that mean roughly the same thing"). - You are building retrieval-augmented generation (RAG) - finding the most relevant documents to feed to an LLM. - You are doing recommendations based on similarity. the code #### cURL ``` curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/vector/shop.products/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "sku-9281", "embedding": [0.0124, -0.0883, 0.0451, /* ... 768 floats ... */], "dim": 768, "metric": "cosine", "metadata": { "category": "running-shoes", "price": 129.0 } }' ``` #### Python ``` # embedding_768d is your 768-float list, e.g. from an OpenAI/Cohere call. db.vector.put( "shop.products", "sku-9281", embedding_768d, metadata={ "category": "running-shoes", "price": 129.0 }, ) ``` #### TypeScript ``` await db.vectorPut("shop.products", { id: "sku-9281", embedding: embedding768d, // number[] of length 768 dim: 768, metric: "cosine", metadata: { category: "running-shoes", price: 129.0 }, }); ``` #### Go ``` err := db.VectorPut(ctx, "shop.products", originchain.VectorPutRequest{ ID: "sku-9281", Embedding: embedding768d, // []float32 of length 768 Dim: 768, Metric: "cosine", Metadata: map[string]any{"category": "running-shoes", "price": 129.0}, }) ``` what each field means | Field | Type | Required | What it is | | --- | --- | --- | --- | | id | string | yes | A unique ID for this vector. Usually the primary key of the row it was extracted from (here, the SKU). | | embedding | float[] | yes | The vector itself. The length must match `dim` exactly. | | dim | int | yes | The vector's length. Must be the same value for every vector in this table - the first insert locks it in. | | metric | string | no | How "closeness" is measured. `cosine` (default) for most text models. `l2` for distance-style models. `dot` for inner-product. Locked after the first insert. | | metadata | object | no | Anything you want to filter on at search time. Example: `{ "category": "shoes" }` lets you later restrict the search to shoes only. | `` `` `common mistakes Wrong dim. If your embeddings are 1536 floats but the first insert said dim: 768, every later insert fails with 400 dim_mismatch. The first insert sets the lock. Mixing metrics. If you start with cosine and later send l2, you get 400 metric_mismatch. Pick one and stay with it. Filtering on un-indexed metadata. Filters work on any key, but they are fastest on simple equality (category == "shoes"). Range filters on price are slower. tip · skip the separate call Vectors are stored on their own endpoint, not as a row column. The id you pass here is what links the vector back to a row (typically the row's primary key). See Vector tables for the full reference.` ``4. Insert a search document. what this does Index a piece of text so it shows up in keyword search. OriginChainDB breaks the text into tokens, applies the analyzer you declared on the schema (lowercase, stemming, etc.), and builds an inverted index that's ranked by BM25 - the standard relevance algorithm used by Elasticsearch and Lucene. when to use it You want users to find rows by typing keywords ("carbon plate marathon shoes"). You want phrase search ("exact phrase in quotes"). You want fuzzy matching that tolerates typos. the code cURL Python TypeScript Go POST /v1/tenants/:t/fts/:table/:field curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/fts/shop.products/description" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "sku-9281", "text": "Lightweight road runner with a carbon plate, designed for marathon pace." }'db.fts.index( "shop.products", "description", doc_id="sku-9281", text="Lightweight road runner with a carbon plate, designed for marathon pace.", )await db.ftsIndex("shop.products", "description", { doc_id: "sku-9281", text: "Lightweight road runner with a carbon plate, designed for marathon pace.", });err := db.FTSIndex(ctx, "shop.products", "description", originchain.FTSIndexRequest{ DocID: "sku-9281", Text: "Lightweight road runner with a carbon plate, designed for marathon pace.", }) what each field means Field Where Required What it is :table URL yes The schema name. Same as in row writes. :field URL yes The column name to index this text under. You can index the same column with different documents. doc_id body yes A unique ID for this document. Usually the row's primary key. Re-indexing with the same doc_id replaces the previous text - no stale matches. text body yes The actual text to index. No size limit on this endpoint, but very large documents (~MBs) are better split into multiple doc_ids. common mistakes Indexing a doc_id that doesn't match a row. The FTS index is independent of the row store - nothing stops you from indexing a doc_id that doesn't exist in the table. You'll get search hits pointing at nothing. Treat doc_id as "the primary key of the row this text describes" and you'll be fine. Wrong analyzer for your language. The default English analyzer doesn't stem German or Chinese well. Declare the right analyzer on the schema (Snowball stemmers in 18 languages plus CJK / Thai / Khmer tokenizers). Indexing the wrong text. The text you put here is what gets searched - if you index only the product name, users can't search by description. Concatenate every searchable field into one string before indexing. tip · skip the separate call Like vectors, full-text indexes live on their own runtime endpoint - they're not declared on the row schema. Re-indexing the same doc_id replaces the previous text in the same write, so there are no stale postings. 5. Insert a graph relationship. what this does Create a link (an edge) between two rows that you can later walk. Examples: a product is supplied by a supplier, a user follows another user, an order belongs to a customer. Here is the important thing: an edge is not a separate write. You declare which columns are relations on the schema, then the engine creates and maintains the forward + reverse edges automatically whenever you write the row. when to use it You want to query things like "all products from supplier X" or "all orders placed by user Y" without writing a JOIN. You want to do multi-hop walks like "friends of friends" or "products bought by users who bought this one". You want shortest-path queries. the code Assuming the schema has [[relations]] column = "supplier_id" target = "shop.suppliers" declared, this row write creates the edge automatically: cURL Python TypeScript Go POST /v1/tenants/:t/rows/:schema · edge written atomically # A graph edge is NOT a separate write. # Declare `[[relations]]` on the schema (see "Try it yourself" below), # then write the row - the engine creates the forward and reverse # edges automatically because `supplier_id` is declared as a relation. curl -X POST "$ORIGINCHAIN_URL/v1/tenants/$T/rows/shop.products" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "sku-9281", "name": "Carbon Marathon", "supplier_id": "sup-44", "price_cents": 12900 }'# Same put as Section 1. The schema declares supplier_id as a # relation pointing at shop.suppliers, so the edge is created # automatically when the row is saved. db.rows.put("shop.products", { "id": "sku-9281", "name": "Carbon Marathon", "supplier_id": "sup-44", "price_cents": 12900, })// Same row write as Section 1 - the edge is implicit. await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/shop.products`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "sku-9281", name: "Carbon Marathon", supplier_id: "sup-44", price_cents: 12900, }), });body, _ := json.Marshal(map[string]any{ "id": "sku-9281", "name": "Carbon Marathon", "supplier_id": "sup-44", "price_cents": 12900, }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/shop.products", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") http.DefaultClient.Do(req) common mistakes Target row doesn't exist yet. If sup-44 doesn't exist in shop.suppliers, the edge is still stored - it just points at a non-existent row. Decide whether you want this. To enforce existence, add a foreign-key constraint on the schema. Forgetting that updates retire the old edge. If a product's supplier_id changes from sup-44 to sup-77, the old edge is removed in the same write. Good for accuracy, surprising if you expected history. Trying to add an edge without a column for it. Edges piggy-back on columns. If you want a many-to-many relationship without a column, create a join table (e.g., shop.product_tags) and put relations on its columns. next · walk the edges Once edges are written, see Graph queries for how to walk them (neighbors, BFS, shortest path, PageRank). 6. All four at once (atomic). what this does Save the same product as a row, a vector embedding, a search document, and a graph edge - in three coordinated calls, all backed by the same record. This is the pattern most real apps end up with. In a typical stack you'd write the row to Postgres, push the embedding to a vector database, push the text to Elasticsearch, and trust three systems to stay in sync. Here, all four projections live in the same instance and commit together atomically. when to use it You are building a product catalog that needs to be searchable by exact filter, by similarity, by keyword, and by relationship - all at once. You are building a RAG pipeline that also needs structured filtering. You want to stop maintaining three separate databases. the schema One schema, all four shapes declared up front. Register this with POST /v1/tenants/$T/schemas: # manifest.toml - the row schema. Defines columns + a graph edge. # Vector and full-text indexes are NOT declared here - they live on # their own runtime endpoints (see /docs/vector, /docs/fts) and link # back to rows by primary key. namespace = "shop" table = "products" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" [[columns]] name = "supplier_id" ty = "str" [[columns]] name = "price_cents" ty = "i64" # money in minor units - never f64 [[columns]] name = "description" ty = "str" # Secondary index on supplier_id so neighbor lookups are fast. [[indexes]] name = "by_supplier" columns = ["supplier_id"] # Turn supplier_id into a graph edge: the row write creates the edge # automatically because [[relations]] is declared. [[relations]] name = "supplied_by" from_col = "supplier_id" bidirectional = true [relations.target] namespace = "shop" table = "suppliers" pk = "id" the code Save one product. Each call lands as one atomic write on the engine - if any of the three calls fails, you can safely retry the failing one because the calls are idempotent. Python TypeScript Go row + vector + full-text + supplier edge # One product, written four ways - but it is ONE write from the # database's perspective. If anything fails, nothing is saved. product_id = "sku-9281" description = "Lightweight road runner with a carbon plate, designed for marathon pace." # 1. The row itself. The graph edge to `shop.suppliers` is created # automatically because `supplier_id` is a declared relation. db.rows.put("shop.products", { "id": product_id, "name": "Carbon Marathon", "supplier_id": "sup-44", "price_cents": 12900, "description": description, }) # 2. The vector embedding. Your app computes the float[]; the engine stores it. db.vector.put( "shop.products", product_id, embed(description), # 768-float list metadata={ "category": "running-shoes", "price": 129.0 }, ) # 3. The BM25 full-text index. Re-indexing the same doc_id replaces # the old postings - no ghost matches. db.fts.index( "shop.products", "description", doc_id=product_id, text=description, )const productId = "sku-9281"; const description = "Lightweight road runner with a carbon plate, designed for marathon pace."; // 1. The row + graph edge (raw fetch until row helpers ship). await fetch(`${BASE_URL}/v1/tenants/${TENANT}/rows/shop.products`, { method: "POST", headers: { "Authorization": `Bearer ${OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: productId, name: "Carbon Marathon", supplier_id: "sup-44", price_cents: 12900, description, }), }); // 2. The vector embedding. await db.vectorPut("shop.products", { id: productId, embedding: await embed(description), // number[] of length 768 dim: 768, metric: "cosine", metadata: { category: "running-shoes", price: 129.0 }, }); // 3. The BM25 full-text index. await db.ftsIndex("shop.products", "description", { doc_id: productId, text: description, });productId := "sku-9281" description := "Lightweight road runner with a carbon plate, designed for marathon pace." // 1. The row + graph edge (raw http until row helpers ship). rowBody, _ := json.Marshal(map[string]any{ "id": productId, "name": "Carbon Marathon", "supplier_id": "sup-44", "price_cents": 12900, "description": description, }) req, _ := http.NewRequestWithContext(ctx, "POST", BASE_URL+"/v1/tenants/"+TENANT+"/rows/shop.products", bytes.NewReader(rowBody)) req.Header.Set("Authorization", "Bearer "+OC_TOKEN) req.Header.Set("Content-Type", "application/json") http.DefaultClient.Do(req) // 2. The vector embedding. db.VectorPut(ctx, "shop.products", originchain.VectorPutRequest{ ID: productId, Embedding: embed(description), // []float32 of length 768 Dim: 768, Metric: "cosine", Metadata: map[string]any{"category": "running-shoes", "price": 129.0}, }) // 3. The BM25 full-text index. db.FTSIndex(ctx, "shop.products", "description", originchain.FTSIndexRequest{ DocID: productId, Text: description, }) what just happened After those three calls, the same product is visible to four kinds of query: SQL: SELECT * FROM shop.products WHERE price < 150 Vector: find the 10 products most similar to a query embedding Full-text: find products whose description matches "marathon carbon" Graph: find all products supplied by sup-44 See Querying your data for each of these. common mistakes Forgetting one of the three calls. Row, vector, and full-text are stored independently. If you insert the row but skip the embedding, the product won't show up in vector search. Wrap the three calls in your own helper function so they always go together. Embedding the wrong text. Embed what you want users to search by - usually a title + description, not the SKU. Ignoring failures. Each call returns success or failure. If only the vector call fails, the row and FTS index are still written. Decide whether you want to roll back manually (delete the row) or just retry the failing call. ← Docs home Next: query your data →`` --- # PostgreSQL ingest connector on OriginChainDB Canonical source: https://originchaindb.com/docs/integrations/postgres-ingest Sitemap last modified: 2026-09-14T15:11:01.000Z docs · integrations · postgres ingest # PostgreSQL ingest connector. The PostgreSQL ingest connector pulls rows from an existing Postgres source into your OriginChainDB tenant in one HTTP call. No separate ETL service, no queue, no glue code. Configure the source, declare the table mappings, POST — the engine handles the rest. ## Quickstart. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/ingest/postgres/sync" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source": { "host": "pg.internal", "port": 5432, "database": "production", "user": "readonly", "password": "***", "sslmode": "require" }, "mappings": [ { "source_table": "public.orders", "target_schema": "orders", "primary_key": ["id"], "columns": ["id", "customer_id", "amount", "status", "created_at"] } ] }' ``` ## Source config. - `host / port / database` Source endpoint. Must be reachable from your OriginChainDB instance — public, peered, or via a private endpoint on Enterprise. - `user / password` Read-only Postgres role recommended. Credentials are held in memory for the duration of the sync. - `sslmode` "require" or "verify-full". Plaintext is rejected. ## Target schema mapping. Each mapping picks a source table, a target OriginChainDB schema, a primary key, and the columns to pull. The connector creates the OriginChainDB schema if it does not exist, with column types inferred from the Postgres type map. You can declare the schema up-front in the manifest if you want explicit control. ## Scheduling. v1 is one-shot: each POST runs a complete sync. Schedule from your side (cron, a queue, or a tiny scheduler script) until the engine ships a managed scheduler. The endpoint is idempotent on primary key — the same row pulled twice overwrites; no duplicates. #### Python ``` import os, requests, schedule, time def sync_once(): requests.post( f"{os.environ['OC_ENDPOINT']}/v1/tenants/{os.environ['T']}/ingest/postgres/sync", headers={"Authorization": f"Bearer {os.environ['OC_TOKEN']}"}, json={ "source": {...}, "mappings": [...], }, timeout=900, ).raise_for_status() schedule.every(15).minutes.do(sync_once) while True: schedule.run_pending() time.sleep(1) ``` ## What's next. - Logical-decoding mode (CDC-style stream) is on the roadmap. v1 is full-table sync. - Source-side column filters (push down predicates to the Postgres query) — coming alongside CDC. - MySQL + SQLite connectors are next on the integration list. --- # Multi-node configurations - OriginChainDB Canonical source: https://originchaindb.com/docs/multi-node Sitemap last modified: 2026-09-13T20:18:50.000Z operate · multi-node # Multi-node configurations A multi-node configuration runs your instance across more than one node while presenting exactly one endpoint. You get more write throughput and bigger rate budgets; your application code, schemas, and queries stay the same. ## What multi-node gives you. Write throughput that scales Adding nodes adds write capacity. Measured: a 2-node configuration sustains ~2x the write throughput of a single node on the same workload. One endpoint Your application talks to a single URL with a single token, whatever the node count. No client-side routing, no per-node connection strings, no code changes when the configuration grows. Automatic data distribution Data is distributed across nodes table by table, automatically. You never assign tables to nodes or manage placement - create tables the way you always have. Atomic cross-node transactions A transaction can span tables on different nodes and still commits atomically - all writes land or none do. Same BEGIN / COMMIT / ROLLBACK surface as a single node. Rate budgets that scale with you API-key budgets multiply by node count. A 3-node configuration delivers 3x the aggregate request, byte, and query budgets - through the same single endpoint. Per-node resilience Each node can optionally carry its own synchronized standby in a separate zone, selected through the resilience configurations - the same options you know from single-node. Throughput and budget numbers here are measured, not projected: ~2x write throughput at 2 nodes on our standard workload. Larger node counts add capacity in the same direction; measure against your own workload before sizing for a hard target. ## One endpoint, no client changes. Every request - reads, writes, vector search, full-text, graph, natural language - goes to the same base URL you'd use on a single node. OriginChainDB routes it to the right node internally, and [transactions](https://originchaindb.com/docs/transactions) that touch tables on different nodes commit atomically. A small set of operations that would have to be evaluated across nodes is refused with a clear error instead of being answered partially - see [current limits](https://originchaindb.com/docs/multi-node#multi-node-limits) below. The same is true of your API keys and [rate budgets](https://originchaindb.com/docs/rate-limits): one key, one endpoint, and the budgets you're entitled to across the whole configuration are delivered through that single endpoint. You don't spread traffic across nodes to collect your full budget - it's aggregate by contract. ## Current limits - refused honestly, never answered wrong. A few operations would need to be evaluated across every node at once. Rather than return a partial or incorrect result, a multi-node configuration refuses these with a clear 501 Not Supported naming the operation, so your application can detect and handle them explicitly. This is a deliberate safety decision: you get an honest error, never silently wrong data. - Cross-node SQL subqueries - a query whose subquery needs rows that live on more than one node. - Cross-node access policies - role-based access control and row-level security that would have to be evaluated across nodes. - Natural-language /ask - these queries fan out across the whole configuration. - Ad-hoc cross-node writes - a single write spanning nodes outside the supported [transaction](https://originchaindb.com/docs/transactions) path. If your workload needs any of these today, a single-node configuration supports all of them with no restrictions - and multi-node is the right choice when your access patterns are node-local (keyed lookups, per-tenant or per-entity queries) and you need horizontal capacity. Full parity for these operations is on the roadmap; until then they are a documented, predictable error. ## How data is distributed. Distribution is by table, and automatic. When you create tables, OriginChainDB places them across the nodes of your configuration; queries and writes are routed accordingly. There is no placement API to learn and no distribution key to design - the table remains the unit you think in. Because distribution is table-by-table, workloads with several independently busy tables see the most immediate benefit: their write load spreads naturally across nodes. ## Choosing a node count. | Your situation | Start with | | --- | --- | | Most applications - a single node's throughput covers you, and you can grow later. | 1 node | | Sustained write volume near a single node's ceiling, or rate budgets that keep hitting 429 despite batching. | 2 nodes | | High-volume ingest across many busy tables, or aggregate budget requirements a smaller configuration can't meet. | 3+ nodes | Two sizing notes. First, batch before you scale: batch endpoints are dramatically more efficient than single-row calls (see [rate limits → batching](https://originchaindb.com/docs/rate-limits#batching)) - many workloads that look like they need more nodes actually need batching. Second, node count is a write-throughput and budget decision; if your bottleneck is a single slow query, look at [query tuning](https://originchaindb.com/docs/query) and indexes first. ## Growing later. Node count is not a one-way door decision. Nodes can be added to a running configuration online, with no downtime - the endpoint stays the same, and data redistributes behind it. The console exposes your configuration's node layout and is where growth happens. Practical consequence: start with the smallest configuration that covers today's measured load, and grow when your own numbers say so. --- # NQL: natural-language queries by industry - OriginChainDB Canonical source: https://originchaindb.com/docs/nql/industries Sitemap last modified: 2026-09-13T20:18:50.000Z nql · by industry # Natural-language querying by industry Natural language is worth it where the questions are many and unpredictable, and the asker cannot write SQL. Each of these pairs the convenience with the guardrail it needs. ## Operations: the question nobody wrote a dashboard for Dashboards answer the questions someone anticipated. Incidents produce the others. ### Ask it directly No ticket to the data team, no new panel, and the answer comes from live rows rather than yesterday's export. ``` { "nl": "which services had more than 50 errors in the last hour?" } ``` show the plan next to the answer An operator acting on a number needs to know which table it came from. Return the plan alongside the rows and the answer becomes checkable. ## Finance: self-service without a warehouse export The recurring monthly questions are known; the follow-ups are not. ### Scope it and ask Naming the schemas keeps the compiler on the tables that hold the truth, and keeps it away from ones it should not touch. ``` { "nl": "settled volume by currency last month", "schemas": ["fin.payments"] } ``` permissions still apply NQL runs the plan as the caller. Row-level security and column masking apply to the generated query exactly as to a hand-written one, so a natural-language interface cannot become a way around them. ## Support: triage in the agent's own words An agent asks about a customer mid-conversation and needs the answer now, not a query builder. ### One sentence, one answer Ordinary phrasing, over the same tables the product writes to. ``` { "nl": "open tickets for acme corp older than three days" } ``` ## Internal tools: an assistant with a real backend The assistant should read the database, not a copy of it, and should be honest when it cannot. ### Plan first, then run Compile without executing, check the plan against what the user is allowed to ask, then run it. That sequence is what makes an assistant safe to point at production. ``` { "nl": "...", "plan_only": true } // inspect { "nl": "..." } // then execute ``` an agent can call this directly The MCP server exposes ask, SQL, vector and search as tools, so an AI IDE or agent can use these surfaces without bespoke glue. ## What these have in common - The plan is the trust boundary. Inspect it and natural language becomes a query surface rather than a guess. - Scope beats cleverness. Naming the schemas is the single biggest improvement in both accuracy and latency. - Your policies still hold. The generated query runs as the caller, under the same row and column rules. [Start from zero Ask, scope, inspect the plan.](https://originchaindb.com/docs/nql/quickstart)[Ask AI The wider AI surface this sits in.](https://originchaindb.com/docs/ai) --- # NQL quickstart - natural-language queries on OriginChainDB Canonical source: https://originchaindb.com/docs/nql/quickstart Sitemap last modified: 2026-09-23T08:33:23.000Z nql · quickstart # Ask your first question in English Ask a question in English and inspect the plan compiled against your registered schemas. ## Before you start An instance, an API key, and at least one registered schema — the compiler can only use tables it knows about. ``` export OC_URL='https://' export OC_TENANT='' export OC_TOKEN='' ``` ## The four steps 1. 1. Ask One field. The answer comes back as rows. ``` curl -X POST "$OC_URL/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"nl":"how many orders were placed last week?"}' ``` 2. 2. Scope it to the right tables Naming the schemas narrows what the compiler considers. Leave it out and every registered schema the acting database user holds SELECT on is a candidate, which is slower because the catalog is larger. ``` { "nl": "revenue by brand", "schemas": ["shop.orders", "shop.products"] } ``` 3. 3. Read the plan before you trust it Ask for the plan alongside the rows. There is no compile-without-executing mode — the question still runs. This is the step that turns a demo into something you can put in front of a user: you can see exactly which tables and predicates it chose. ``` { "nl": "top customers by spend", "show_plan": true } ``` 4. 4. Let the repeats be cheap The same question asked again reuses its compiled plan rather than recompiling. See [the reference](https://originchaindb.com/docs/ask) for how the cache is keyed and invalidated. what it will not do The compiler answers from your schemas. It does not invent tables, and a question it cannot ground is refused rather than answered with a plausible guess — which is the failure mode that makes natural-language querying untrustworthy elsewhere. ## Examples Every operation below is shown in cURL, Python, TypeScript and Go. SDK coverage varies more here than anywhere else on the docs — each tab says what it can and cannot do. 1 ## Ask a question. One required field: `nl`. Note the name — it is not `question` or `query`. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "the 5 largest paid orders", "schemas": ["shop.orders"] }' ``` #### Python ``` result = db.ask("the 5 largest paid orders", schemas=["shop.orders"]) for row in result["rows"]: print(row["id"], row["amount_cents"]) print(result["cache"]) # "hit" or "miss" ``` #### TypeScript ``` const result = await db.ask("the 5 largest paid orders", { schemas: ["shop.orders"], }); for (const row of result.rows) console.log(row); console.log(result.cache); // "hit" or "miss" ``` #### Go ``` // Go's Ask takes only the question - it sends no schemas allowlist // and cannot request the plan. Call the endpoint directly if you need those. resp, err := db.Ask(ctx, "the 5 largest paid orders") if err != nil { return err } fmt.Println(resp.Cache, len(resp.Rows)) ``` response ``` { "rows": [ { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "amount_cents": 48200, "status": "paid", "placed_ms": 1714478049000 }, { "id": "01JTRX9KQ3YH8K2WMX0F5JZAC1", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ004", "amount_cents": 31150, "status": "paid", "placed_ms": 1714477012000 } ], "cache": "miss" } ``` `rows` is the answer. `cache` tells you whether the question was answered from a previously compiled plan (`"hit"`) or compiled fresh (`"miss"`) — a cache hit skips the model round trip entirely. what the response does not contain There is no confidence score, no token usage, no model name, and no generated SQL string. You get rows, a cache flag, and — on request — the plan. If you were hoping to gate on a confidence value before showing results to a user, that signal does not exist; validate by inspecting the [plan](https://originchaindb.com/docs/nql/quickstart#plan) instead. the cache is keyed on your question and your schema Questions are normalised before caching — whitespace collapsed, case folded outside quoted strings, trailing punctuation stripped — so "Show me 5 paid orders" and "show me 5 paid orders." share a cache entry. The key also folds in a hash of every registered manifest, so changing a schema automatically invalidates every cached plan. You never have to clear it by hand after a migration. 2 ## Seeing the compiled plan. `show_plan: true` returns the compiled plan alongside the rows. This is the single most useful thing on this page — it is how you check that the question was understood the way you meant it. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "the 5 largest paid orders", "schemas": ["shop.orders"], "show_plan": true }' ``` #### Python ``` # The Python SDK's ask() does not expose show_plan - call the # endpoint directly when you need the compiled plan back. import httpx r = httpx.post( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/ask", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "nl": "the 5 largest paid orders", "schemas": ["shop.orders"], "show_plan": True, }, ) print(r.json()["plan"]) ``` #### TypeScript ``` // TypeScript is the only SDK that exposes show_plan. const result = await db.ask("the 5 largest paid orders", { schemas: ["shop.orders"], show_plan: true, }); console.log(JSON.stringify(result.plan, null, 2)); ``` #### Go ``` // The Go SDK cannot send show_plan - its AskResponse.Plan field // therefore stays empty. Use net/http for the plan. body, _ := json.Marshal(map[string]any{ "nl": "the 5 largest paid orders", "schemas": []string{"shop.orders"}, "show_plan": true, }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/ask", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` response ``` { "rows": [ /* … */ ], "cache": "miss", "plan": { "op": "limit", "n": 5, "child": { "op": "sort", "keys": [ { "path": "amount_cents", "order": "desc" } ], "child": { "op": "filter", "predicate": { "op": "eq", "path": "status", "value": "paid" }, "child": { "op": "scan", "schema": "shop.orders" } } } } } ``` Read it inside-out: scan `shop.orders`, keep rows where `status = "paid"`, sort by `amount_cents` descending, take 5. That is exactly the question, and you can see it is exactly the question — which is the whole point. - The plan returned is the post-optimisation one, so it may differ slightly from what the compiler first emitted. - Adding `?explain` to the URL also turns the plan on, and additionally returns an `explain` field with a cost-annotated tree. Careful: any value enables it — `?explain=false` switches it on just as surely as `?explain=true`. - TypeScript is the only SDK that exposes `show_plan`. Python's `ask()` has no such parameter; Go's `Ask()` takes only the question, so its `AskResponse.Plan` field can never be populated. 3 ## What it compiles to. Not to SQL. A question compiles to the same internal query plan that a SQL statement compiles to — the structure you saw above. The executor cannot tell which surface produced it, which is why NL inherits the same optimiser and the same execution guarantees as everything else. There is no SQL string anywhere in the pipeline, so there is nothing to show you if you were looking for "the SQL it wrote". The plan is the compiled query. ### Two compilers, in order. This is worth knowing because it explains a lot of the behaviour: 1. A deterministic rule compiler runs first. It understands a compact grammar — roughly `[top N] [where ] [by ] [limit N]` — with the operators `=`, `==`, `is`, `!=`, `<>`, `>` and `<`, plus `in` for value lists. If your question fits, you get a plan with no model call at all — fast, free and perfectly repeatable. 2. Only if the rule compiler cannot parse the question does it fall through to the managed language model, which emits a plan in the same JSON shape. Phrasing a question in the rule grammar's shape — "top 5 shop.orders where status = paid by amount_cents desc" — is therefore a real performance technique, not a stylistic one. Note the rule compiler has no `>=` or `<=`, cannot mix `and` with `or` in one condition, and does no aggregation — those questions go to the model. ### What the plan grammar can express. | Supported | Not available through NL | | --- | --- | | Table scans and column projections Index lookups Filters — `eq`, `ne`, `gt`, `lt`, `in`, `and`, `or` Sorting and limits Grouping with `count`, `sum`, `avg`, `min`, `max` Single-hop relation walks Catalog questions ("what tables are there") | Joins — the grammar has no join shape Any other aggregate (median, percentile, stddev, count-distinct) `DISTINCT`, window functions, CTEs, subqueries Set operations (`UNION` and friends) Graph algorithms and variable-length paths Any write — insert, update, delete Any schema change | ask is read-only, structurally A question cannot modify data, whatever it says. The endpoint executes through the read path, which holds an immutable handle on storage and refuses write operations outright — "delete all cancelled orders" returns a `400`, not a deletion. There is no schema-mutation shape in the grammar at all, so DDL is not expressible in the first place. NL also does not participate in transactions. joins are the big one The engine supports joins perfectly well from SQL — but the NL plan grammar has no join shape, so a question spanning two tables cannot compile into one. Ask "which customers spent the most" across `orders` and `customers` and you will get an error or an answer drawn from one table only. Single-hop relation walks are the closest available thing. For genuine multi-table questions, write the [SQL](https://originchaindb.com/docs/schemas/sql). 4 ## Grounding it in your schema. Omit `schemas` and every table the acting database user holds SELECT on becomes context. Pass it and only the listed tables are visible. On an instance with more than a handful of tables, scoping is the highest-leverage thing you can do for both accuracy and latency. #### cURL ``` # Unscoped - every schema registered on the instance becomes context. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "how many orders are unpaid" }' # Scoped - only these tables are visible to the compiler. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/ask" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nl": "how many orders are unpaid", "schemas": ["shop.orders"] }' ``` #### Python ``` # Unscoped - every registered schema becomes context. db.ask("how many orders are unpaid") # Scoped - only these tables are visible to the compiler. db.ask("how many orders are unpaid", schemas=["shop.orders"]) ``` #### TypeScript ``` // Unscoped - every registered schema becomes context. await db.ask("how many orders are unpaid"); // Scoped - only these tables are visible to the compiler. await db.ask("how many orders are unpaid", { schemas: ["shop.orders"], }); ``` #### Go ``` // Go's Ask() sends no schemas allowlist, so it always runs unscoped. // Call the endpoint directly to scope it. body, _ := json.Marshal(map[string]any{ "nl": "how many orders are unpaid", "schemas": []string{"shop.orders"}, }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/ask", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` - Ids must be fully qualified — `shop.orders`, not `orders`. An id that is not registered is a `400`. - Only names and types are sent — never any of your row data. No sample rows are included in the compiler's context. - There is no size cap on that context, so an instance with very many tables produces a very large prompt. Scope it. - Both Python and TypeScript expose `schemas`. Go does not — its `Ask()` always runs unscoped. table names are validated, column names are not Every table a plan references is checked against your catalog, and a plan naming an unknown table is rejected before it runs. Columns get no such check. A plan that filters on a column you do not have will pass validation and be executed — typically returning zero rows rather than an error. An empty answer to a question you expected to match something is the signature of this, and `show_plan` is how you confirm it. 5 ## Failure modes. Compilation problems are all `400`s with a `compile:` prefix. The useful thing is that they are specific — the engine distinguishes "I don't understand this" from "this is ambiguous" from "that table doesn't exist". ### An ambiguous question. If a bare table name matches two namespaces, you get a hard error rather than a guess — and notably, the question is not passed to the model to be resolved. Ambiguity is treated as your intent to name a table, stated imprecisely. ``` { "error": "compile: ambiguous schema \"orders\" — qualify as .
" } ``` ### A question it cannot compile. An unsupported aggregation — "the median order value" — or a question the model cannot turn into a valid plan ends here. The model is given up to three attempts, each one shown the previous rejection so it can self-correct; if all three fail you get: ``` { "error": "compile: unparseable natural-language query: …" } ``` | Situation | Status | | --- | --- | | Empty or whitespace-only question | 400 | | Ambiguous or unknown table | 400 | | Ambiguous graph traversal direction | 400 | | Unsupported aggregation or shape | 400 | | A question that implies a write | 400 | | No schemas registered on the instance | 400 | | Answer exceeds the result budget | 413 | | Rate or concurrency cap hit | 429 | | Monthly credit exhausted | 402 | the quiet failure is the dangerous one Every case above is loud. The failure to actually watch for is the plausible wrong answer: a question that compiles cleanly into a plan that is not quite what you meant — a dropped qualifier, a filter on the wrong column, a limit that was silently ignored. It returns `200` and looks fine. This is inherent to the approach, not a bug, and it is exactly why `show_plan` exists. Check the plan before you trust a number. a compile can take a while A cache miss means a model round trip, with a 60-second ceiling per attempt and up to three attempts — so a pathological question can sit for a couple of minutes before returning a `400`. Set a client timeout you are comfortable with, and don't put an uncached NL call in a request path with a tight budget. 6 ## Availability and caps. not available on a free database Natural language is a paid-database feature. On a free database the console does not show the NL option at all — the query workbench's language switcher reads SQL / Cypher / Search — and the endpoint is gated. Select a paid instance and the option reappears. On a paid instance, three independent limits apply. Because each compile is genuinely expensive, `/ask` is metered separately from ordinary database calls — running out of NL credit never affects the rest of your database. | Limit | Window | On breach | | --- | --- | --- | | Ask credits | Monthly | 402 | | Concurrent calls in flight | Instantaneous | 429 | | Request rate | Per second | 429 | Both the concurrency and per-second allowances scale with the configuration you are running — the entry configuration allows 5 questions in flight at once, the standard configuration 15, and the advanced configuration 50. There is no per-day cap; the windows are monthly, instantaneous and per-second, and nothing else. 402 — credits exhausted ``` { "error": "quota_exhausted", "resource": "ask", "used": 5000, "limit": 5000, "message": "monthly /ask quota exhausted (5000 / 5000). top up credits at https://app.originchain.ai/billing.", "top_up_url": "https://app.originchain.ai/billing" } ``` 429 — too many in flight ``` { "error": "concurrent_cap_exceeded", "resource": "ask", "in_flight": 5, "cap": 5, "configuration": "entry", "msg": "Tenant has 5 concurrent ask call(s) in flight; the entry configuration allows 5. Wait for in-flight calls to complete or upgrade your configuration." } ``` The concurrency rejection carries `Retry-After: 1` and is genuinely transient — retry it. The `402` carries no `Retry-After`, because retrying will not help until credits are topped up. Live in-flight usage is visible on `GET /v1/tenants/:tenant/usage`. There is no maximum question length beyond the 8 MB request-body limit — but a long question is not a better one, and everything past the part that names tables and conditions is noise the compiler has to work through. 7 ## NL vs writing the query yourself. | Reach for NL when… | Write SQL when… | | --- | --- | | You are exploring unfamiliar data A non-technical user is asking the question The question is one-off and the cost of being slightly wrong is low You want a starting point to refine into real SQL | The query runs more than once It sits on a latency-sensitive path Correctness is not negotiable It needs a join, a set operation, or an aggregate outside the five basics It writes anything at all | The most productive pattern is to treat NL as a drafting tool: ask the question with `show_plan: true`, read the plan to learn what the engine thinks your data means, then write the equivalent SQL and ship that. You get the exploration speed without putting a compile step in your production path. For a user-facing "ask your data" feature, cache aggressively at your own layer too. Identical questions already hit the engine's plan cache, but the questions real users type are rarely byte-identical. ## What to read next [NQL reference The request shape, the plan format and the cache.](https://originchaindb.com/docs/ask)[Help the compiler Schema hints that make questions resolve the way you expect.](https://originchaindb.com/docs/schemas/nl)[By industry Ops dashboards, revenue questions, support triage.](https://originchaindb.com/docs/nql/industries)[Grounded answers When you want prose with citations rather than rows.](https://originchaindb.com/docs/rag) --- # Ops and runbook - OriginChainDB managed database Canonical source: https://originchaindb.com/docs/ops Sitemap last modified: 2026-09-23T08:33:23.000Z 07 · ops & runbook # Ops - what to alert on, how to fail over, how to recover. Monitor instance health and follow the documented procedures for backups, recovery, failover, and schema changes. ## Health checks | endpoint | what it tells you | alert when | | --- | --- | --- | | /health | Liveness - process is up and the store is mountable. Returns 200 + JSON build/version. Use for load-balancer health checks. | any 5xx or sustained timeout | | /ready | Readiness - engine has caught up to the latest durable state and the standby (if any) is connected. | non-200 for >30s after boot | | /v1/tenants/:t/usage | Per-tenant usage snapshot: row count, vector count, in-flight queries, subscription state. | watch via dashboard - poll for trends | ### Per-tenant usage signal `GET /v1/tenants/:t/usage` returns the current snapshot: row count per schema, vector count per table, in-flight Ask queries, subscription status. Poll it from your monitoring system; the dashboard renders the same data live. A first-class Prometheus exporter is on the roadmap. Until then, hit `/usage` on a poll interval and graph the deltas. ## Write durability A write is acknowledged only after it has been flushed to durable storage on the instance. The flush has to return before the response is sent, not after - so a `2xx` means your data is on disk rather than sitting in memory, and it survives the database process being killed. This holds for row writes, SQL data changes, Cypher writes, transaction commits, vector writes and sequence values. Secondary index structures are flushed on a short timer instead. Full-text index writes, materialized-view installs and refreshes, and geospatial index updates are written to durable storage but are not individually flushed before the response - they ride a periodic flush. A crash can therefore lose the last few of them even though the response said `2xx`. Your underlying rows are unaffected: re-running the index or refresh operation rebuilds these from row data, which is why they are treated as derived state rather than a source of record. Two recovery caveats are worth knowing, because both are about getting data back after the flush rather than about the flush itself. We publish them because they change what you should do operationally. - Abrupt power loss - a hard reset can interrupt a write mid-flight and leave a partially-written record behind. Your acknowledged data is still on disk, but the instance does not always reopen unattended: the incomplete record has to be cleared first. In our power-cut testing this occurred in 2 of 6 hard resets, and after trimming every acknowledged write was present - we measured no data loss. Plan for this as an availability event that can need an assisted restart, not as data loss. - A full storage volume - while the volume is full, writes fail with a `5xx`. Nothing is silently dropped and that part is safe. Once space is reclaimed the instance starts accepting writes again and serves them on read, but those writes are not reliably durable until the instance has been restarted and checked - in our testing they did not survive the next restart. Treat a full-volume alert as an incident: reclaim space, then restart and verify before you trust writes made in that window. Alert on free space well before it runs out. A fix for both is complete in the engine and has not yet reached running instances. This page describes the build your instance is running today, and we will narrow it once the fix has rolled out. Separately, durability on the instance is not the same as durability across a failover - replication to the standby is asynchronous, see [failover](https://originchaindb.com/docs/ops#failover) below. ## Backups Dedicated configurations have three layers, all archived in the same region as the tenant. Free and Starter do not include point-in-time recovery. - Daily storage snapshot - always on, and the floor for every dedicated configuration. A crash-consistent snapshot of the instance's storage lands in an encrypted vault once a day, so worst-case exposure without the layers below is the time since the last snapshot. - Off-box archive shipping - opt-in, preview, and not currently something to plan around. Recent changes ship to the archive on a size-and-time trigger rather than per commit. In practice the recoverable point has been advancing far more slowly than that design implies - measured across the fleet it sits around half an hour behind, and on some instances the archive has not advanced at all. We are fixing it. Until we say otherwise, size your recovery expectations from the daily snapshot above, not from this layer. - Base snapshot shipping - a fresh base is shipped to the archive every 30 minutes. Restore picks the latest base at-or-before the target and replays the archive up to it. - Restore-to-timestamp - choose a wall-clock target (RFC 3339) from the console or API and OriginChainDB rebuilds a fresh instance from the chosen snapshot + archive. The result opens cleanly through the same validating loaders that serve live traffic. Recovery points are retained 30 days; manual purges on data-subject requests are documented in the runbook. ## Continuous backup streaming Opt-in via the Intra-Segment PITR add-on The default flow gives PITR granularity at the archive cadence - fine for most tenants, too coarse for compliance-heavy ones. The continuous backup stream is designed to tighten it: the tail ships on a timer whenever enough has changed since the last ship. It does not make the archive per-commit - below the change threshold nothing ships until the next base, so a quiet instance has a coarse recovery point by design. It is also not currently delivering that design: on the instances we measured, the tail has stopped advancing and the recoverable point sits roughly half an hour behind, sometimes further. Treat sub-second or per-second recovery as something we have not delivered rather than something you can switch on, and talk to us before you build a compliance commitment on this layer. On the managed service the continuous backup stream runs as a built-in task alongside restore-side replay (auto-replay of the latest tail at-or-before the target). ## Failover When the primary stops renewing its writer lease - a wedged process, or a host that has gone silent - the in-region standby is the instance that takes over as primary. Automatic promotion is opt-in and off by default: the standby has to be configured for it, and the switch into writer mode is completed by a hook outside the database process, so check which path your instance is configured for. Either way the standby cannot claim the lease until the old primary's claim has expired and a grace window has passed. A standby that is configured to promote itself refuses to do so while it is behind the primary - by default it has to be fully caught up, so a lagging standby stays a standby until an operator decides otherwise. Replication to the standby is asynchronous, so a promotion after an abrupt primary loss can cost the most recent acknowledged writes. What happens during a failover: 1. Fence the old primary - the old primary loses the writer lease and refuses further writes with a 503 for as long as another instance holds that lease. 2. Promote the standby - the in-region standby takes over as the new primary and begins serving reads and writes. 3. Health-check - the new primary is verified healthy on `/health` before traffic is cut over. 4. Update DNS - your endpoint is re-pointed at the new primary with a 60s TTL; propagation is typically sub-minute. Failover is initiated on a confirmed dead primary - a genuine host failure, not a transient network blip. Your application keeps the same endpoint; once DNS propagates, writes resume against the promoted instance. When failover is held off: the primary is responding but slow (we load-shed rather than fail over); the replica is far behind on replication (failing over would lose data - replication is investigated first); or an online schema migration is mid-backfill. ## Schema migrations Online migrations are first-class - no read-only window, no service bounce. The model: - Rate-limited backfill - the backfill is capped at a fraction of live throughput so production traffic stays prioritised. - Dual-read transform - readers see the new (v1) shape during backfill via an on-the-fly transform applied to old (v0) rows. - Atomic cutover - the version bump is a single commit; reads switch to the v1-native shape on the next read. - Abort before cutover - once cutover lands, the only path forward is an inverse-rewrite migration. Aborts during backfill are safe and reversible. Use online migrations any time the manifest version bumps. Troubleshooting a migration stuck mid-backfill is in [incident response](https://originchaindb.com/docs/ops#incident). ## Observability - EXPLAIN - prefix any SELECT with `EXPLAIN` to return the plan tree without running the query. Useful for verifying that your indexes are being used. See [SQL reference → EXPLAIN](https://originchaindb.com/docs/sql#explain). - Per-tenant /usage - `GET /v1/tenants/:t/usage` returns row counts, vector counts, in-flight queries, and subscription state. Poll it for monitoring. - OTLP push tracing - opt-in. Configure the collector endpoint via the dashboard; spans are pushed for every `/v1` request with tail-based sampling that retains slow + error traces in full. - Dashboard - live metrics tiles for every instance: query latency, replication lag, recent writes, error rate. Visit [app.originchain.ai](https://app.originchain.ai/). ## Incident response Our internal incident playbook covers the full procedure. Highlights: - Pager severity: sev 1 (tenant down) → 5 min ack, 1 hr fix-or-mitigate. Sev 2 (one alarm tripped) → 15 min / 4 hr. Sev 3 (drift) → next business day. - Status page: publishes per-region health and incident timelines. Subscribers get email + webhook on any sev 1. - Migration stuck mid-backfill: abort the migration via `POST /v1/tenants/:t/migrations/:id/abort`, then resubmit. - Order of operations: stop the bleeding, find root cause, write a postmortem, fix the underlying problem. Quiet incidents become loud ones. Incidents today are handled by core engineering during extended business hours with best-effort overnight coverage; the pager-severity SLAs above are the targets we hold ourselves to. 24/7 named-engineer coverage is available on Enterprise - contact sales. ## Compliance posture - SOC 2 Type 1: underway with Vanta/Drata and an external CPA. Contact for audit timeline; the in-flight gap analysis is available to procurement under NDA. - HIPAA BAA: available on Enterprise. PHI workloads must run in a region the BAA covers and on a dedicated-capacity instance. - GDPR DPA: available on Enterprise. EU-region instances support the DPA by default; data-subject deletion follows our documented runbook. --- # Query OriginChainDB - SQL, vector, full-text, graph, NL Canonical source: https://originchaindb.com/docs/query Sitemap last modified: 2026-09-23T08:33:23.000Z how-to · query # Query data Choose SQL, vector, graph, full-text, or natural-language queries to retrieve the data your application needs. Each shape gets its own reference page below. The [HTTP API reference](https://originchaindb.com/docs/api) lists every endpoint in one place. SDKs ([Python, TypeScript, Go](https://originchaindb.com/docs/sdk)) wrap each one. [SQL → Use when you need filters, GROUP BY, aggregates, or JOINs. SELECT with WHERE / GROUP BY / LIMIT plus INNER / LEFT / RIGHT / FULL OUTER joins (up to 32 tables).](https://originchaindb.com/docs/sql)[Vector search → Use for semantic search, RAG, and recommendations. Top-k nearest-neighbor lookup over embeddings with optional metadata filtering.](https://originchaindb.com/docs/vector)[Full-text → Use for keyword search. BM25 ranked, boolean AND, or exact-phrase. 18-language stemming and lemmatization in 9.](https://originchaindb.com/docs/fts)[Graph → Use for relationship walks - neighbors, multi-hop BFS, shortest path, PageRank. 19 algorithms.](https://originchaindb.com/docs/graph)[Natural language (Ask) → Use when you'd rather write English than SQL. Sentence in, rows out. Compiled plans are cached so repeat questions are fast.](https://originchaindb.com/docs/ask) ## Shared mechanics. Every shape compiles to a JSON-serialisable Plan tree, cached by question hash, replayable. `?explain=true` on any read endpoint returns the executed plan annotated with per-node row counts and µs timings. Cancel an in-flight plan with the ULID handed back in `X-OC-Query-Id`. The operator catalog is Scan, IndexScan, Filter, HashJoin, OuterJoin, Aggregate, Sort, Limit, and RelationHop. --- # Quickstart - Create a table and query it Canonical source: https://originchaindb.com/docs/quickstart Sitemap last modified: 2026-09-24T06:54:16.000Z 01 · quickstart # Create your first table and query Create a table, insert sample data, and run SQL, vector, full-text, and graph queries against one OriginChainDB instance. Every code example is shown in cURL, Python, TypeScript, and Go. Pick the tab you want and the rest of the page follows that language. 1 ## Sign up and get your token. Head to [app.originchain.ai/signup](https://app.originchaindb.com/signup). No card is required to start: the Free configuration gives you 1 GB and the full engine at $0. You add a card only when you pick a paid configuration. After signup you will: 1. Pick a region close to your users. Your data never leaves it. 2. Pick a configuration (you can change it later). Start with the smallest one. 3. Wait for the live instance status to confirm that provisioning has completed. 4. Copy three things from the "Connect" panel: your endpoint URL, a one-time bearer token, and your tenant ID. Open the connection guideOriginChainDB console [[Image: Connection guide open over the current instance overview, with the application header and sidebar visible.]](https://originchaindb.com/_astro/connect.DDybhNZM.png) The current connection guide. Use the live endpoint and credentials issued for your instance.Current console · Dark theme · Local preview save them to your shell ``` # Paste these into your shell. Replace the values with the ones # from your dashboard after you finish step 1. export OC_HOST=acme.db.originchain.ai export OC_TENANT=acme export OC_TOKEN=oc_live_xxxxxxxxxxxxxxxx ``` one-time token The bearer token is shown once and never again. Paste it into your secrets manager (1Password, Doppler, etc.) before you close the tab. If you lose it, rotate it from the dashboard and update your apps. 2 ## Install an SDK (or skip if you prefer cURL). The SDKs handle authentication, retries, and request signing for you. You can use raw HTTP too - every example below shows cURL alongside. install #### cURL ``` # cURL is built into macOS and most Linux distros - check it's there: curl --version # On Windows: comes with Git Bash and PowerShell 6+. ``` #### Python ``` pip install originchain ``` #### TypeScript ``` npm install @originchain/sdk ``` #### Go ``` go get github.com/originchain-ai/originchain-go ``` connect #### cURL ``` # cURL has no client - every call carries the bearer header. # Confirm your env vars resolve and the instance is reachable: curl -s -o /dev/null -w "HTTP %{http_code}\n" \ "https://$OC_HOST/health" \ -H "Authorization: Bearer $OC_TOKEN" # → HTTP 200 if everything is wired up. ``` #### Python ``` import os from originchain import OriginChain db = OriginChain( base_url=f"https://{os.environ['OC_HOST']}", bearer=os.environ["OC_TOKEN"], tenant=os.environ["OC_TENANT"], ) # `db` is your handle for the rest of this page. ``` #### TypeScript ``` import { OriginChainClient } from "@originchain/sdk"; const db = new OriginChainClient({ baseUrl: `https://${process.env.OC_HOST}`, bearer: process.env.OC_TOKEN!, }); // `db` is your handle for the rest of this page. ``` #### Go ``` package main import ( "context" "os" "github.com/originchain-ai/originchain-go" ) var ( ctx = context.Background() db = originchain.NewClient(originchain.Config{ BaseURL: "https://" + os.Getenv("OC_HOST"), Bearer: os.Getenv("OC_TOKEN"), }) ) // `db` and `ctx` are your handles for the rest of this page. ``` SDK coverage today Python wraps every endpoint on this page. TypeScript and Go wrap SQL, vector, full-text, graph, and the ask endpoint - row reads and writes use raw `fetch` / `net/http` for now (helpers ship in the next release). 3 ## Define your first table. A schema is the shape of your data - the list of columns, their types, and any extra indexes or extractions you want. OriginChainDB reads schemas as small TOML files. Save this as `schemas/orders.toml`. It declares an `orders` table that you can later query four ways - SQL, vector, full-text, graph. schemas/orders.toml ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" # ULIDs / UUIDs travel as text required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" # money in minor units - never f64 [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # epoch milliseconds # Make WHERE status = ... fast. [[indexes]] name = "by_status" columns = ["status"] ``` what each block does | Block | What it does | | --- | --- | | namespace, table | A two-part name for your table. You address it as `namespace.table` in queries (here, `shop.orders`). | | primary_key | A list of column names that uniquely identify a row. Most tables use a single ID column - put it in the array. | | [[columns]] | One column per block. Set `required = true` on columns that can't be null. Six types: `str`, `i64`, `u64`, `f64`, `bool`, `bytes`. | | [[indexes]] | Secondary index for fast filtering. Without this, `WHERE status = ...` scans every row. | Notice there's no column type for vectors or full-text indexes. Those live on separate runtime endpoints - you'll see how in steps 7 and 8. The schema TOML is just for rows. See [Schema reference](https://originchaindb.com/docs/schemas) for relations, foreign keys, CHECK constraints, and derived columns. register it #### cURL ``` # Save the TOML above as schemas/orders.toml, then: curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/schemas" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: text/plain" \ --data-binary @schemas/orders.toml ``` #### Python ``` with open("schemas/orders.toml") as f: db.schemas.register(f.read()) ``` #### TypeScript ``` import { readFileSync } from "node:fs"; const toml = readFileSync("schemas/orders.toml", "utf-8"); await db.registerSchema(toml); ``` #### Go ``` tomlBytes, _ := os.ReadFile("schemas/orders.toml") req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/schemas", bytes.NewReader(tomlBytes)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "text/plain") http.DefaultClient.Do(req) ``` common mistakes - Wrong Content-Type. TOML schemas use `text/plain`, not `application/json`. - No primary key. Every schema needs a top-level `primary_key = ["..."]` array naming at least one declared column, otherwise registration fails. 4 ## Save your first row. One call saves the row and updates every secondary index that touches its columns - all atomically, in a single durable commit. (Vector and full-text indexes are managed separately - we'll do those in steps 7 and 8.) #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "amount_cents": 12950, "status": "paid", "notes": "rush delivery, signed by recipient", "placed_ms": 1714478049000 }' ``` #### Python ``` db.rows.put("shop.orders", { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "amount_cents": 12950, "status": "paid", "notes": "rush delivery, signed by recipient", "placed_ms": 1714478049000, }) ``` #### TypeScript ``` // TS SDK doesn't wrap row writes yet - using fetch directly. await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/shop.orders`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "01JTRX9KQ3YH8K2WMX0F5JZAB7", customer: "01JTRX1H4Q9P0N2WMX0F5JZ001", amount_cents: 12950, status: "paid", notes: "rush delivery, signed by recipient", placed_ms: 1714478049000, }), }, ); ``` #### Go ``` // Go SDK doesn't wrap row writes yet - using net/http directly. body, _ := json.Marshal(map[string]any{ "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "amount_cents": 12950, "status": "paid", "notes": "rush delivery, signed by recipient", "placed_ms": uint64(1714478049000), }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/rows/shop.orders", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") http.DefaultClient.Do(req) ``` For bulk imports (1k+ rows), see [Insert → bulk](https://originchaindb.com/docs/insert#bulk). The endpoint takes JSON arrays or NDJSON streams. 5 ## Read it back. Look up a row by its primary key. This is the fastest read in OriginChainDB - direct hash lookup, single-digit milliseconds. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/01JTRX9KQ3YH8K2WMX0F5JZAB7" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` row = db.rows.get("shop.orders", "01JTRX9KQ3YH8K2WMX0F5JZAB7") print(row["status"], row["amount_cents"]) # → paid 12950 ``` #### TypeScript ``` // TS SDK doesn't wrap row reads yet. const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/shop.orders/01JTRX9KQ3YH8K2WMX0F5JZAB7`, { headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}` } }, ); const row = await res.json(); console.log(row.status, row.amount_cents); ``` #### Go ``` req, _ := http.NewRequestWithContext(ctx, "GET", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/rows/shop.orders/01JTRX9KQ3YH8K2WMX0F5JZAB7", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, _ := http.DefaultClient.Do(req) defer resp.Body.Close() var row map[string]any json.NewDecoder(resp.Body).Decode(&row) fmt.Println(row["status"], row["amount_cents"]) ``` 6 ## Run a SQL query. Filter, group, and aggregate with regular SQL. This query finds every paying customer and totals their orders. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT customer, SUM(amount_cents) AS total FROM shop.orders WHERE status = '\''paid'\'' GROUP BY customer LIMIT 10" }' ``` #### Python ``` result = db.sql(""" SELECT customer, SUM(amount_cents) AS total FROM shop.orders WHERE status = 'paid' GROUP BY customer LIMIT 10 """) # result.kind == "select"; rows are dicts. for row in result.rows: print(row["customer"], row["total"]) ``` #### TypeScript ``` const result = await db.sql(` SELECT customer, SUM(amount_cents) AS total FROM shop.orders WHERE status = 'paid' GROUP BY customer LIMIT 10 `); if (result.kind === "select") { for (const row of result.rows) { console.log(row.customer, row.total); } } ``` #### Go ``` result, err := db.SQL(ctx, ` SELECT customer, SUM(amount_cents) AS total FROM shop.orders WHERE status = 'paid' GROUP BY customer LIMIT 10 `) if err != nil { /* handle */ } if result.Kind == "select" { for _, row := range result.Rows { fmt.Println(row["customer"], row["total"]) } } ``` The console's Workbench provides a SQL editor and results viewer. Its local preview uses a sample dataset and a limited read-only parser; run the engine examples on this page against your configured live instance. Run a SQL queryOriginChainDB console [[Image: Current OriginChainDB console with Workbench selected in the sidebar, the SQL editor, and local sample query results.]](https://originchaindb.com/_astro/sql.33U5scMy.png) The current Workbench running sample SQL; the API examples on this page require the configured shop dataset.Current console · Dark theme · Local preview what you can use today The SQL surface supports SELECT, WHERE, GROUP BY, aggregates (COUNT, SUM, AVG, MIN, MAX), INNER / LEFT / RIGHT / FULL OUTER JOIN (up to 32 tables), LIMIT, ORDER BY, HAVING, ranking window functions, CTEs (`WITH` / `WITH RECURSIVE`), correlated subqueries and UNION/INTERSECT/EXCEPT, plus INSERT, UPDATE, DELETE, and transactions. Not yet supported: `NATURAL JOIN`, `?` bind-param placeholders (use `$1`), and a few join-shaped combinations - `HAVING` or a `DISTINCT` aggregate across a JOIN, and a window function in the same `SELECT` as a JOIN. See [SQL reference](https://originchaindb.com/docs/sql) for the full status. 7 ## Vector search. Vector data lives on its own endpoint, separate from rows. You store one embedding per row (linked by the row's primary key) and then run top-k similarity search against the query embedding. save the embedding (cURL only - SDK examples below) #### cURL ``` # Save an embedding for the order. dim + metric are set per-call; # the first put locks them in for the table. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "embedding": [0.12, -0.04, 0.91, 0.33, -0.18, 0.05, 0.42, -0.27], "dim": 8, "metric": "cosine" }' ``` search #### cURL ``` # Find the 10 orders whose embeddings are closest to a query vector. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.10, -0.05, 0.88, 0.30, -0.20, 0.07, 0.40, -0.25], "k": 10, "dim": 8, "metric": "cosine" }' ``` #### Python ``` # Put first, then query. The "id" links the vector back to the row. db.vector.put( "shop.orders", "01JTRX9KQ3YH8K2WMX0F5JZAB7", [0.12, -0.04, 0.91, 0.33, -0.18, 0.05, 0.42, -0.27], ) hits = db.vector_topk( "shop.orders", query=[0.10, -0.05, 0.88, 0.30, -0.20, 0.07, 0.40, -0.25], k=10, dim=8, metric="cosine", ) for hit in hits: print(hit.id, hit.score) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.orders", { query: [0.10, -0.05, 0.88, 0.30, -0.20, 0.07, 0.40, -0.25], k: 10, dim: 8, metric: "cosine", }); for (const hit of hits) { console.log(hit.id, hit.score); } ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.orders", originchain.VectorTopKRequest{ Query: []float32{0.10, -0.05, 0.88, 0.30, -0.20, 0.07, 0.40, -0.25}, K: 10, Dim: 8, Metric: "cosine", }) if err != nil { /* handle */ } for _, h := range hits { fmt.Println(h.ID, h.Score) } ``` Recall and latency are tunable. See [Vector reference](https://originchaindb.com/docs/vector) for index choice (HNSW vs IVF), `mode=fast` vs `mode=high_recall`, and metadata filtering. 8 ## Full-text search. Same pattern as vectors: full-text data lives on its own endpoint. You index the text once (linked to the row's primary key as `doc_id`), then search. `mode=bm25` ranks hits by relevance (the same algorithm Elasticsearch and Lucene use). index the text (cURL only - SDK examples below) #### cURL ``` # Index the text under (table, field, doc_id). doc_id = the row's PK. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "text": "rush delivery, signed by recipient" }' ``` search #### cURL ``` # Then search. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes?q=rush+delivery&mode=bm25&k=10" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` # Index + search. db.fts.index( "shop.orders", "notes", doc_id="01JTRX9KQ3YH8K2WMX0F5JZAB7", text="rush delivery, signed by recipient", ) hits = db.fts.search( "shop.orders", "notes", q="rush delivery", mode="bm25", k=10, ) for hit in hits: print(hit.score, hit.doc_id) ``` #### TypeScript ``` await db.ftsIndex("shop.orders", "notes", { doc_id: "01JTRX9KQ3YH8K2WMX0F5JZAB7", text: "rush delivery, signed by recipient", }); const hits = await db.ftsSearch("shop.orders", "notes", { q: "rush delivery", mode: "bm25", k: 10, }); for (const hit of hits) { if (typeof hit === "string") console.log(hit); else console.log(hit.score, hit.doc_id); } ``` #### Go ``` db.FTSIndex(ctx, "shop.orders", "notes", originchain.FTSIndexRequest{ DocID: "01JTRX9KQ3YH8K2WMX0F5JZAB7", Text: "rush delivery, signed by recipient", }) hits, err := db.FTSSearch(ctx, "shop.orders", "notes", originchain.FTSSearchRequest{ Q: "rush delivery", Mode: "bm25", K: 10, }) if err != nil { /* handle */ } for _, h := range hits { fmt.Println(h.Score, h.DocID) } ``` Other modes: `mode=boolean` (every term must match, no ranking), `mode=phrase` (exact phrase). See [Full-text reference](https://originchaindb.com/docs/fts) for fuzzy search, highlights, and facets. 9 ## Walk a graph. The `customer` column links every order to a customer. Any column you'd later use as a graph edge needs a `[[relations]]` block on the schema (we'll add one in [Insert → graph](https://originchaindb.com/docs/insert#graph)). Then walking is one call: #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/graph/shop.orders/neighbors?rel=customer&pk=01JTRX1H4Q9P0N2WMX0F5JZ001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` neighbors = db.graph.neighbors( "shop.orders", rel="customer", pk="01JTRX1H4Q9P0N2WMX0F5JZ001", ) print(len(neighbors), "orders for this customer") ``` #### TypeScript ``` const neighbors = await db.graph.neighbors("shop.orders", { rel: "customer", pk: "01JTRX1H4Q9P0N2WMX0F5JZ001", }); console.log(neighbors.length, "orders for this customer"); ``` #### Go ``` neighbors, err := db.Graph().Neighbors(ctx, "shop.orders", originchain.NeighborsRequest{ Rel: "customer", PK: "01JTRX1H4Q9P0N2WMX0F5JZ001", }) if err != nil { /* handle */ } fmt.Println(len(neighbors), "orders for this customer") ``` More than one hop? See [Graph reference](https://originchaindb.com/docs/graph) for BFS, shortest path, PageRank, and community detection. ## Where to go next. [schemas Schema reference Every column type, every option, every extraction.](https://originchaindb.com/docs/schemas) [insert Bulk inserts + atomic multi-shape writes Streaming millions of rows, writing rows+vectors+text+edges in one go.](https://originchaindb.com/docs/insert) [query Pick the right query shape When to use SQL vs vector vs full-text vs graph vs ask.](https://originchaindb.com/docs/query) --- # RAG on OriginChainDB - retrieval-augmented generation Canonical source: https://originchaindb.com/docs/rag Sitemap last modified: 2026-09-23T08:33:23.000Z use case · rag # RAG / LLM apps Retrieve relevant chunks from OriginChainDB and pass them to a language model as context for an answer. The big advantage over a Pinecone + Elasticsearch + Postgres stack: rows, vector embeddings, and full-text indexes all live in one place. No ETL, no eventual consistency between systems, one bearer token. Your app's authorization rules apply to retrieval automatically because there's only one store. ## 1. Schema for chunks. A chunk is one passage of text addressable as a single row. Each chunk gets a primary key (typically `docId:chunkIndex`), the source document ID for filtering, page / location metadata for citations, and the chunk text itself. ``` # manifest.toml - the row schema for your chunks. # Vector embeddings and the BM25 index live on separate runtime endpoints, # linked back to this row by primary key. namespace = "rag" table = "chunks" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "doc_id" ty = "str" required = true [[columns]] name = "source" ty = "str" [[columns]] name = "page" ty = "i64" [[columns]] name = "text" ty = "str" [[columns]] name = "created_ms" ty = "u64" # Index doc_id so "all chunks for this doc" is fast. [[indexes]] name = "by_doc" columns = ["doc_id"] ``` Register this with `POST /v1/tenants/:t/schemas` (see [Schemas overview](https://originchaindb.com/docs/schemas)). Vector dim and FTS analyzer are set per-call at the runtime endpoints - not on this schema. ## 2. Ingest path. For each chunk: one row write, one vector put, one FTS index call. Three separate endpoints today (helpers that combine them ship in a future release). Each call is atomic on its own. #### Python ``` # 1. Chunk the source text into ~512-token windows with 50-token overlap. # 2. Embed each chunk with whatever embedding model you use. # 3. Write the chunk row + the embedding + the FTS index entry. from originchain import OriginChain from openai import OpenAI import time, os db = OriginChain( base_url=f"https://{os.environ['OC_HOST']}", bearer=os.environ["OC_TOKEN"], tenant=os.environ["OC_TENANT"], ) ai = OpenAI() def ingest(doc_id, source, full_text): chunks = chunk_text(full_text, tokens=512, overlap=50) emb = ai.embeddings.create( model="text-embedding-3-small", input=[c.text for c in chunks], ) for i, c in enumerate(chunks): cid = f"{doc_id}:{i}" # 1. The row. db.rows.put("rag.chunks", { "id": cid, "doc_id": doc_id, "source": source, "page": c.page, "text": c.text, "created_ms": int(time.time() * 1000), }) # 2. The vector embedding. db.vector.put("rag.chunks", cid, emb.data[i].embedding, metadata={"source": source, "page": c.page}) # 3. The BM25 index. db.fts.index("rag.chunks", "text", doc_id=cid, text=c.text) ``` #### TypeScript ``` // TS SDK doesn't wrap row writes yet - using fetch for the row write, // SDK for vector + FTS. Same logic as Python above. import { OriginChainClient } from "@originchain/sdk"; import OpenAI from "openai"; const db = new OriginChainClient({ baseUrl: `https://${process.env.OC_HOST}`, bearer: process.env.OC_TOKEN!, }); const ai = new OpenAI(); async function ingest(docId: string, source: string, fullText: string) { const chunks = chunkText(fullText, { tokens: 512, overlap: 50 }); const emb = await ai.embeddings.create({ model: "text-embedding-3-small", input: chunks.map(c => c.text), }); for (let i = 0; i < chunks.length; i++) { const cid = `${docId}:${i}`; // 1. Row write via fetch. await fetch(`https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/rag.chunks`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: cid, doc_id: docId, source, page: chunks[i].page, text: chunks[i].text, created_ms: Date.now(), }), }); // 2. Vector embedding. await db.vectorPut("rag.chunks", { id: cid, embedding: emb.data[i].embedding, dim: 1536, metric: "cosine", metadata: { source, page: chunks[i].page }, }); // 3. BM25 index. await db.ftsIndex("rag.chunks", "text", { doc_id: cid, text: chunks[i].text }); } } ``` common mistakes - Chunks too big. Most embedding models lose quality past ~512 tokens. Past ~1500 you're throwing away precision. - No overlap. Without 10-20% chunk overlap, key sentences land at chunk boundaries and disappear from retrieval. - Embedding only - or BM25 only. Vector misses exact tokens (SKUs, error codes, model numbers). BM25 misses synonyms. Index both, fuse the results. ## 3. Retrieval + fusion. Run vector and BM25 in parallel, fuse the results, fetch the full chunks for the top-N fused IDs, send to the LLM with a strict "use only this context" instruction. #### Python ``` # 1. Embed the question. # 2. Run vector topk + BM25 search in parallel. # 3. Fuse with Reciprocal Rank Fusion. # 4. Pull the chunk rows for the top-k IDs. # 5. Send them to the LLM as context. def answer(question, source_filter=None): # 1. q_vec = ai.embeddings.create( model="text-embedding-3-small", input=question, ).data[0].embedding # 2. Parallel retrieval (one round-trip each, can run concurrently). vec_hits = db.vector_topk( "rag.chunks", query=q_vec, k=20, dim=1536, metric="cosine", filter={"source": source_filter} if source_filter else None, ) fts_hits = db.fts.search("rag.chunks", "text", q=question, mode="bm25", k=20) # 3. Reciprocal Rank Fusion. Pick the top 5 fused IDs. fused_ids = rrf([h.id for h in vec_hits], [h.doc_id for h in fts_hits])[:5] # 4. Pull full chunk text + metadata. rows = [] for cid in fused_ids: rows.append(db.rows.get("rag.chunks", cid)) # 5. Send to the LLM as bounded, cited context. context = "\n\n".join( f"[Source {i+1}] ({r['source']} p.{r['page']}) {r['text']}" for i, r in enumerate(rows) ) completion = ai.chat.completions.create( model="gpt-4o", temperature=0.0, messages=[ { "role": "system", "content": "Answer using ONLY the context below. Cite every claim with [Source N]. " "If the context is insufficient, say so." }, { "role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}" }, ], ) return completion.choices[0].message.content ``` #### TypeScript ``` async function answer(question: string, sourceFilter?: string) { // 1. Embed the question. const qVec = (await ai.embeddings.create({ model: "text-embedding-3-small", input: question, })).data[0].embedding; // 2. Parallel retrieval. const [vecHits, ftsHits] = await Promise.all([ db.vectorTopk("rag.chunks", { query: qVec, k: 20, dim: 1536, metric: "cosine", filter: sourceFilter ? { source: sourceFilter } : undefined, }), db.ftsSearch("rag.chunks", "text", { q: question, mode: "bm25", k: 20 }), ]); // 3. Reciprocal Rank Fusion. Pick top 5. const fusedIds = rrf( vecHits.map(h => h.id), (ftsHits as { doc_id: string }[]).map(h => h.doc_id), ).slice(0, 5); // 4. Fetch chunks via raw fetch (row helpers ship soon). const rows = await Promise.all(fusedIds.map(async (cid) => { const r = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/rag.chunks/${cid}`, { headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}` } }, ); return r.json(); })); // 5. Send to LLM. const context = rows.map((r, i) => `[Source ${i+1}] (${r.source} p.${r.page}) ${r.text}`, ).join("\n\n"); const completion = await ai.chat.completions.create({ model: "gpt-4o", temperature: 0.0, messages: [ { role: "system", content: "Answer using ONLY the context below. Cite every claim with [Source N]. " + "If the context is insufficient, say so." }, { role: "user", content: `Context:\n${context}\n\nQuestion: ${question}` }, ], }); return completion.choices[0].message.content; } ``` ## 4. The RRF helper. Reciprocal Rank Fusion. Five lines. Works for any number of ranked lists. `k=60` is the constant from the original paper - rarely worth tuning. ``` # Reciprocal Rank Fusion - simple, language-agnostic. # Takes ranked lists of IDs, returns a fused ranked list. # k=60 is the standard constant from the original paper. def rrf(*ranked_lists, k=60): scores = {} for lst in ranked_lists: for rank, doc_id in enumerate(lst): scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank + 1) return [doc_id for doc_id, _ in sorted(scores.items(), key=lambda x: -x[1])] ``` ## 5. Tips + tradeoffs. - Optional reranker. A cross-encoder reranker (Cohere Rerank, Voyage Rerank, or a self-hosted bge-reranker) on the fused top-20 typically gains 5-15% NDCG. Add it as step 4.5. - Filter on metadata for tenant isolation. Pass `filter: {tenant_id: "..."}` on vector topk so users only see their own chunks. The filter applies during search, not after - so it's still fast. - Mode selector for speed. Set `mode: "fast"` on the vector topk when latency matters more than recall (e.g., search-as-you-type). The hit count drops slightly but the call gets ~3x faster. - Semantic cache for hot questions. Embed user questions into a small `rag.query_cache` table. Before going through retrieval, do a vector topk on it - if any past question has cosine ≥ 0.97 to this one, return the cached answer. - Cite every claim. Always pass page numbers + source IDs in the context. Force the LLM to cite. Users trust answers they can verify. --- # Rate limits - OriginChainDB per-key request budgets Canonical source: https://originchaindb.com/docs/rate-limits Sitemap last modified: 2026-09-13T20:18:50.000Z reference · rate limits # Rate limits Every OriginChainDB API key carries four independent budgets - requests per second, bytes per second, natural-language queries per second, and concurrent queries - and they scale with your configuration. They exist to bound the damage from a runaway loop, a forgotten `while True` or an autoscaling event, not to gate normal usage. ## What's budgeted. Each API key carries four independent budgets. A request is admitted only if it fits all of them; the first budget you exhaust is the one named in the `X-OC-Limit-Hit` header on the 429. | Budget | Counted as | Notes | | --- | --- | --- | | requests / sec | API calls per second | The primary budget. One batch insert of 1,000 rows counts as one request - see [batching](https://originchaindb.com/docs/rate-limits#batching) below. | | bytes / sec | Payload volume per second | Bounds total data movement independently of request count, so a few huge payloads can't starve the instance. | | NL queries / sec | Natural-language (`/ask`) calls per second | Budgeted separately because a cold compile is far more expensive than an ordinary query. | | concurrent queries | Simultaneous in-flight queries | Caps open requests at any instant, independent of per-second rates. | Budgets are per-key, not per-instance. If you create multiple keys (e.g., one per service), each gets its own budget. The figures for your instance depend on your configuration - and they [scale with node count](https://originchaindb.com/docs/rate-limits#multi-node). If you need more headroom, contact [support@originchain.ai](mailto:support@originchain.ai) - production accounts can have budgets raised on request. ## Budgets scale with node count. On a [multi-node configuration](https://originchaindb.com/docs/multi-node), every budget multiplies by the node count: a 3-node configuration carries 3x the requests/sec, bytes/sec, NL queries/sec, and concurrent-query budgets of a single node. You still talk to one endpoint, and that endpoint delivers the full aggregate budget. There is no traffic to split on the client side and no per-node quota to balance - send your traffic to the same URL and the whole contract is available there. ## Batch before you scale. Rate budgets are counted per request, and batch endpoints move many rows per request - which makes batching the single cheapest lever you have. Measured on our standard ingest workload, the batch path delivers roughly 18x the throughput of the same data sent as single-row calls. If you're hitting 429s on an ingest path, switch to [batch inserts](https://originchaindb.com/docs/insert#bulk) before you reach for a bigger configuration. Most "we need more rate limit" conversations end with a batch endpoint. ## Response headers. Every successful response carries your current usage. Use these to back off proactively before you hit the limit. ``` # Every API response includes these headers: X-RateLimit-Limit: 1000 # request budget for this key X-RateLimit-Remaining: 847 # what you have left in the current window X-RateLimit-Reset: 1714478400 # epoch seconds when the window resets # On 429 responses you also get: Retry-After: 12 # seconds to wait before retrying X-OC-Limit-Hit: bytes_per_sec # WHICH budget you exceeded ``` When `X-RateLimit-Remaining` drops near zero, your code can slow itself down voluntarily instead of waiting for the 429 response. ## Handling 429. When you exceed a budget, OriginChainDB returns `429 rate_limited` with a `Retry-After` header (seconds) and an `X-OC-Limit-Hit` header naming the budget you exhausted - so you know whether to slow down, shrink payloads, or cap concurrency. The SDKs handle the retry automatically - they read `Retry-After`, wait, and retry up to 3 times. If you're calling the API directly, do the same: #### Python ``` # The Python SDK handles 429 + Retry-After automatically. # This is what it does under the hood: import time def call_with_retry(fn, max_retries=3): for attempt in range(max_retries + 1): try: return fn() except OCRateLimitedError as e: if attempt == max_retries: raise time.sleep(e.retry_after or 1.0) ``` #### TypeScript ``` // Same logic for raw fetch users: async function callWithRetry(req: () => Promise, maxRetries = 3) { for (let attempt = 0; attempt <= maxRetries; attempt++) { const res = await req(); if (res.status !== 429) return res; if (attempt === maxRetries) return res; const wait = parseInt(res.headers.get("Retry-After") ?? "1") * 1000; await new Promise(r => setTimeout(r, wait)); } throw new Error("unreachable"); } ``` #### Go ``` // Same logic with stdlib http: for attempt := 0; attempt <= 3; attempt++ { resp, err := http.DefaultClient.Do(req) if err != nil { return err } if resp.StatusCode != 429 { return nil } if attempt == 3 { return fmt.Errorf("rate limited") } wait := 1 if h := resp.Header.Get("Retry-After"); h != "" { fmt.Sscanf(h, "%d", &wait) } time.Sleep(time.Duration(wait) * time.Second) } ``` common mistakes - Tight retry loops without backoff. Retrying instantly on 429 just consumes the next minute's quota. Always honor `Retry-After`. - Sharing one key across many workers. Budgets are per-key, so a fleet of workers sharing one key shares one concurrent-query budget. Issue one key per worker or per service. - Bulk inserts in single-row mode. Sending 10,000 single-row inserts will hit the rate limit. Use `_batch` instead - one request, thousands of rows, ~18x the measured throughput. See [batching above](https://originchaindb.com/docs/rate-limits#batching) and [Insert → bulk](https://originchaindb.com/docs/insert#bulk). --- # Feature releases by surface - OriginChainDB Canonical source: https://originchaindb.com/docs/releases Sitemap last modified: 2026-09-13T20:18:50.000Z reference · feature releases # Feature releases by surface. One place per query surface to see what the engine supports today, what is in preview, and what changed. The support lists are read from the engine rather than written by hand, and the date they were read is recorded below. ## How to read this page. Two sources feed it, and they answer different questions. They are kept apart on every section below. - Current state comes from the engine's own capability endpoint, which publishes a support list per capability family. Several of those lists are generated from the same check the engine runs when it accepts or refuses a configuration, and the engine's tests hold others to what the query surface actually does, failing in both directions: a variant that is advertised but refused fails the build, and so does a variant left on the roadmap that the engine in fact serves. The endpoint carries no dates. - History comes from the [changelog](https://originchaindb.com/changelog), which is the only dated record. Each item below links to its entry. Where the changelog has no dated entry for something, it is not listed rather than given an approximate date. The three columns in each table use the engine's own definitions: `supported` is implemented and generally available; `preview` is shipped but deliberately narrow and not yet generally available; `roadmap` is accepted by the schema but refused by the engine at configuration time, so it is not usable today. A dash means the arm is empty. how current this is Support lists were read from the engine's capability endpoint on 2026-09-06. The endpoint publishes 25 capability families; 21 of them describe a query surface and are listed below. The rest describe the deployment - the instruction set of the host CPU, the write topology, the replication mode, and row-level concurrency control - and are not properties of any one surface, so they are not shown here. ## Elasticsearch. A search API shaped like the Elasticsearch REST and Query DSL surface, answered by the same engine as everything else. [Elasticsearch overview](https://originchaindb.com/docs/elasticsearch). no machine-checked matrix for this surface The engine's capability endpoint publishes no family for Elasticsearch, so there is no support list here that the engine's tests hold to account. Rather than write one by hand, this section carries only the dated history below and points you at the reference pages, which document the surface as it is built. ### What changed, and when. The [changelog](https://originchaindb.com/changelog) records no dated entry for Elasticsearch. Until one exists there is no date to quote, so none is given here. The [elasticsearch overview](https://originchaindb.com/docs/elasticsearch) documents what the surface does today. ## SQL. The /sql endpoint and the wire protocols that ride it. This is the surface the capability endpoint covers in most detail. [SQL reference](https://originchaindb.com/docs/sql). | Capability family | Supported today | In preview | On the roadmap | | --- | --- | --- | --- | | Joins sql_joins | — | - inner_equi - outer - expression_projection - expression_aggregate_argument - computed_group_key - cross | - natural · v1.x - cross_composed_shapes · v1.x - non_equi_on · v1.x - disjunctive_on · v1.x - having_over_grouped_join · v1.x - distinct_aggregate_over_join · v1.x - window_with_join · v1.x - order_by_qualified_over_grouped_join · v1.x | | Common table expressions sql_cte | — | - plain - recursive | - chained_recursive · v1.x - recursive_union_distinct · v1.x - recursive_aggregate_projection · v1.x - cte_column_list · v1.x | | Window functions sql_window_functions | — | - ranking - offset - aggregates_over - frame_clauses - named_windows - distribution | - range_value_offsets · v1.x - composed_shapes · v1.x | | Constraints sql_constraint_kinds | - primary_key - unique - foreign_key - check | — | - foreign_key_cascade · v1.x - foreign_key_set_default · v1.x | | INSERT ... ON CONFLICT sql_insert_on_conflict | — | - values_pk_or_unique_target | — | | Materialized view refresh sql_materialized_view_refresh_modes | - on_demand | - incremental_filter - incremental_aggregate | - incremental_aggregate_default_on · v1.x | | Isolation levels sql_isolation_levels | - read_committed (default) | — | - snapshot_isolation · preview - serializable · preview | | Triggers sql_trigger_kinds | - before_after | — | - instead_of · v1.x Phase C | | Stored procedure languages sql_procedure_languages | — | - sql_only - plpgsql_lite | - wasm · v1.x Phase C 4-6 months | | User-defined function languages sql_udf_languages | — | - sql | - wasm · v1.x Phase C 4-6 months | | PostgreSQL wire compatibility postgres_wire | — | - wire_protocol - isolation - auth - tls - types - error_mapping | - binary_bind_params · v1.x - copy_protocol · v1.x - jsonb_column_type · v1.x - multi_dimensional_arrays · v1.x | | MySQL wire compatibility mysql_wire_dialect | — | - ast_translator - protocol | - prepared_statement_projection_typing · v1.x | | TDS wire compatibility mssql_tds_dialect | — | - tsql_ast_translator - isolation - tls | - commercial_driver_interop_verification · v1.x | ### What changed, and when. - [v1.2](https://originchaindb.com/changelog#v1-2-2026-06-08) 2026-06-08 Materialized views with on-demand refresh. Foreign-key and CHECK constraints, with NoAction, Restrict and SetNull as the on-delete actions. The join cap moved from 5 tables to 32. - [v1.0](https://originchaindb.com/changelog#v1-0-2026-05-02) 2026-05-02 GROUP BY with COUNT, SUM, AVG, MIN and MAX; INNER and LEFT, RIGHT and FULL OUTER joins; HAVING filters. ## Graph. Relationship traversal and graph embeddings. The capability endpoint publishes families for the embedding side only. [Graph reference](https://originchaindb.com/docs/graph). | Capability family | Supported today | In preview | On the roadmap | | --- | --- | --- | --- | | Embedding algorithms graph_embedding_algos | - node2vec - graphsage | — | — | | Embedding aggregators graph_embedding_aggregators | - mean | - max_pool - lstm | - trained_aggregators · v1.x | ### What changed, and when. - [v1.2](https://originchaindb.com/changelog#v1-2-2026-06-08) 2026-06-08 GraphSAGE attribute-aware node embeddings, with Mean, MaxPool and LSTM aggregators. Cypher v3 closed CALL subqueries, nested FOREACH, list comprehensions, MERGE, DELETE, CREATE relation and multi-hop undirected matches. - [v1.0](https://originchaindb.com/changelog#v1-0-2026-05-02) 2026-05-02 Forward and reverse one-hop neighbours, breadth-first search to a configurable maximum depth, path reachability, and weighted shortest path by Dijkstra with a caller-supplied weight function. ## Vector. Approximate nearest-neighbour search: which index kinds exist, how vectors are compressed, and where an index lives. [Vector reference](https://originchaindb.com/docs/vector). | Capability family | Supported today | In preview | On the roadmap | | --- | --- | --- | --- | | Index kinds index_kinds | - hnsw - ivf - ivf_pq | — | — | | Quantization quantization | - none - scalar - binary - pq | — | — | | Index residency index_residency | - ram - disk | — | - hybrid · Q4 2026 | ### What changed, and when. - [v1.2](https://originchaindb.com/changelog#v1-2-2026-06-08) 2026-06-08 IVF and IVF-PQ indexes landed end to end. Binary quantization and PQ quantization attach to HNSW or IVF as separate index kinds, so the recall and memory trade-off is chosen per collection. - [v1.0](https://originchaindb.com/changelog#v1-0-2026-05-02) 2026-05-02 HNSW with cosine, dot and L2 metrics, filtered topk by metadata equality, and f32 SIMD distance kernels. ## Natural language. The /ask endpoint, which compiles a natural-language question into a query plan. [Ask reference](https://originchaindb.com/docs/ask). no machine-checked matrix for this surface The engine's capability endpoint publishes no family for Natural language, so there is no support list here that the engine's tests hold to account. Rather than write one by hand, this section carries only the dated history below and points you at the reference pages, which document the surface as it is built. ### What changed, and when. - [v0.10](https://originchaindb.com/changelog#v0-10-2026-04-22) 2026-04-22 Every /ask call compiles in the same region as the tenant's instance, with credentials scoped per instance and a scope locked to a single model identifier. Repeated calls over the same catalog pay the cached-read rate on the schema prompt. - [v0.7](https://originchaindb.com/changelog#v0-7-2026-02-12) 2026-02-12 The plan cache spills to disk, so a cold start reloads hot plans instead of recompiling them. - [v0.5](https://originchaindb.com/changelog#v0-5-2025-12-18) 2025-12-18 The natural-language to query-plan compiler shipped in the first release, alongside HTTP ingress and the row-keyed store. ## Full-text. Ranked text search. The capability endpoint publishes the per-field analysis knobs: tokenizer, lemmatizer and geospatial index kind. [Full-text reference](https://originchaindb.com/docs/fts). | Capability family | Supported today | In preview | On the roadmap | | --- | --- | --- | --- | | Tokenizers fts_tokenizer_kinds | - uax29 - icu | — | — | | Lemmatizers fts_lemmatizers | - none - dictionary_english - dictionary_spanish - dictionary_french - dictionary_german - dictionary_italian - dictionary_portuguese - dictionary_russian - dictionary_dutch - dictionary_swedish | — | — | | Geospatial index kinds geo_index_kinds | - none - r_tree | — | - s2 · Q1 2027 | ### What changed, and when. - [v1.2](https://originchaindb.com/changelog#v1-2-2026-06-08) 2026-06-08 Lemmatization for nine languages - English, Spanish, French, German, Italian, Portuguese, Russian, Dutch and Swedish - reducing inflected forms to a dictionary-backed lemma. ICU and geo tokenizers shipped alongside. - [v1.0](https://originchaindb.com/changelog#v1-0-2026-05-02) 2026-05-02 BM25 ranking at the Lucene defaults, phrase queries by position-list intersection, the UAX #29 Unicode tokenizer, and Snowball stemming for 18 languages. ## What this page does not cover. The gaps are worth stating, because an absence here means an absence in the sources, not a judgement about the surface. - Graph traversal and graph algorithms. The capability endpoint publishes families for graph embeddings only. Traversal, path finding and the algorithm set are documented on the [graph reference](https://originchaindb.com/docs/graph), and the dated entries above cover what the changelog recorded for them. - Query-side full-text features. The endpoint publishes the per-field analysis knobs - tokenizer, lemmatizer, geospatial index kind - not the query operators. Those are on the [full-text reference](https://originchaindb.com/docs/fts). - Per-family detail. Each family carries engineering notes and stated limitations that this page does not reproduce, because a paraphrase of them here is exactly the hand-maintained text this page exists to replace. Read the capability endpoint on your own instance for the full text. - Anything undated. A capability can appear as supported above with no entry under "what changed" - that means the changelog has no dated record of it, not that it landed recently. PostgreSQL, MySQL, Elasticsearch and other product names are trademarks of their respective owners. Naming them here describes wire-protocol and API compatibility only, and implies no affiliation with or endorsement by those projects. --- # Multi-region active-active (in development) on OriginChainDB Canonical source: https://originchaindb.com/docs/replication/multi-region Sitemap last modified: 2026-09-13T20:18:50.000Z docs · replication · active-active # Multi-region active-active (in development). IN DEVELOPMENT — NOT YET AVAILABLE Multi-region active-active is not available for provisioning today. Asking to run a database in this mode is refused when the configuration is installed — it is never accepted and then quietly degraded — and no shipped engine commits any write through consensus. The consensus machinery runs only on engines explicitly configured for engineering drills, the multi-node group this mode needs is still unbuilt, and no date is scheduled. Active-passive replication is the production path for every tenant. This page describes the intended design so you can plan against it; tell us if you need it and we will factor it into how we sequence the work. Active-active is designed to accept writes in every region you run in and commit them in a single, globally consistent order. Each region would serve both reads and writes locally, with the cluster keeping them in agreement. You would trade a little extra write latency for a stronger durability guarantee on every commit. Everything below describes the intended design for this mode, not behaviour you can buy today, and none of it is scheduled. ## What it is designed to do. - Writes in every region. Your application would write to its nearest region and read its own writes — no single "write region" to route around. Today there is exactly one writer: every database runs the active-passive path. - Stronger durability, on the surfaces it would cover. A committed write would be acknowledged only once a majority of the cluster holds it durably, so an acknowledged write survives the loss of a region. No shipped engine commits a write this way today. The design covers row writes, SQL data changes, Cypher writes, transaction commits and schema registration. It does not yet cover vector index writes, full-text index writes, materialized views, sequence values or geospatial indexes — those are node-local, and closing that gap is a precondition for general availability. It would also cost a little more write latency than single-region active-passive. - Failover without moving a write region (roadmap, not built). The intent is that if a region becomes unavailable, the remaining regions keep accepting reads and writes, with no write region to move. No shipped engine does that — it needs the multi-node consensus group described above. Active-passive failover today has a different shape: a provisioned standby tails the writer's stream and can be promoted, and an opt-in automatic promotion path exists. It is off by default, it relies on an external restart hook installed alongside the engine, and it refuses to promote a standby that is not fully caught up. Plan against that shape, not this one. - Consistent reads. Reads are designed to reflect the most recently committed state of the cluster, so clients would see a single, coherent view of the data. ## When to choose it. Active-active will be the right choice when writes need to be accepted in more than one region — a globally distributed application where users in different regions all write — or when you want the strongest durability guarantee on every commit and can accept slightly higher write latency in exchange. If your workload is regional, or sensitive to write latency, the default active-passive setup is faster on the hot path and keeps every byte in the single region you pick. Every database in service runs that path today. ## Getting set up. - There is nothing to switch on, and no date is scheduled. A request to run a database in this mode is refused when the configuration is installed, and the mode needs a multi-node group per database — not something a running database can be switched into today. If it opens, provisioning and tuning would happen with our team: you pick your regions and durability target, and we configure the cluster for your topology. - Before it can carry customer writes, every write surface has to be either replicated across the cluster or explicitly refused; the surfaces listed above are not there yet. - If multi-region writes are on your roadmap, tell us now — it helps us sequence the work, and we will be straight with you about timing. --- # Schemas on OriginChainDB - the shape of your data Canonical source: https://originchaindb.com/docs/schemas Sitemap last modified: 2026-09-23T08:33:23.000Z reference · schemas # Schemas Define a table, its primary key, columns, indexes, and relationships in a TOML schema manifest. Three pages, three reading paths. This page is the overview - read it first. [Tutorial](https://originchaindb.com/docs/schemas/tutorial) walks you through building one from scratch, step by step. [Reference](https://originchaindb.com/docs/schemas/reference) is the field-by-field manual you reach for once you know what you're looking for. ## Three ways to define a schema. Same table, three entry points - they all converge on one registered schema, so use whichever fits how you work. TOML is the canonical form and what the rest of this page teaches; SQL and the dashboard produce the exact same result. [manifest.toml TOML file Declarative and version-controllable - reviewable in a pull request. The canonical form.](https://originchaindb.com/docs/schemas/reference) [CREATE TABLE SQL DDL Familiar Postgres-style DDL over `/sql`. `ALTER` and `DROP` too.](https://originchaindb.com/docs/schemas/sql) [point & click Dashboard The visual Schema designer - add a table, name columns, pick a primary key, Save.](https://originchaindb.com/docs/dashboard/schema) the same table, in code TOML ``` namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "email" ty = "str" [[columns]] name = "signup_ms" ty = "u64" ``` SQL - POST /v1/tenants/:t/sql ``` CREATE TABLE shop.customers ( id TEXT NOT NULL PRIMARY KEY, email TEXT, signup_ms BIGINT ); ``` Prefer to click? The [Schema designer](https://originchaindb.com/docs/dashboard/schema) in the dashboard builds the same manifest for you - no file to write. The rest of this page uses TOML; see the [SQL page](https://originchaindb.com/docs/schemas/sql) for the full DDL surface. ## The smallest possible schema. This is the minimum viable schema - one table with three columns. Save it as `manifest.toml`. ``` # manifest.toml - the smallest schema you can write. namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "email" ty = "str" [[columns]] name = "signup_ms" ty = "u64" ``` what each piece means | Piece | What it does | | --- | --- | | namespace | A logical grouping for related tables. Like a Postgres schema. Use one per product area: `shop`, `analytics`, `auth`. | | table | The table's name. You query it as `namespace.table` - here, `shop.customers`. | | primary_key | A list of column names that uniquely identify a row. Most tables use a single ID column; put it in the array. | | [[columns]] | One block per column. `name` = the column name, `ty` = the type. Set `required = true` on columns that can't be null. | ## Six core types - and purpose-built ones on top. Six base scalars cover most data. On top of them OriginChainDB ships purpose-built types - `timestamp`, `date`, `uuid`, `decimal`, `enum`, `json`, `list`, `inet`, and more - each format-validated on write. See the [schema reference](https://originchaindb.com/docs/schemas/reference) for the full list. | Type | Use it for | | --- | --- | | str | Any text. Names, emails, IDs (including ULIDs and UUIDs - they travel as their canonical text form), category labels, status strings, JSON blobs you don't need to query into. | | i64 | Signed 64-bit integers. Use for money stored in minor units (cents, paise, satoshis), quantities, refundable counters. Never use floats for money. | | u64 | Unsigned 64-bit integers. Monotonic counters, sequence numbers. For event time prefer the native `timestamp` type. | | f64 | Double-precision float. Use for rates, ratios, percentages, scientific measurements - things where rounding doesn't matter. | | bool | A single yes/no value. | | bytes | Opaque binary blob. Use for compressed payloads, encoded vectors stored alongside rows, anything that's already a byte sequence. | ### Purpose-built types Declared the same way (`ty = "timestamp"`) and validated on write - a wrong-shaped value is rejected with a `400`. Each stores on the same substrate as its backing scalar, so filters and range scans stay fast. | Type | For | | --- | --- | | timestamp | Event time - epoch milliseconds. Range-scans like an integer. | | date | Calendar date (epoch days), no time component. | | interval | A duration in milliseconds. | | uuid | RFC-4122 UUID, format-validated. | | decimal | Exact money - e.g. `"19.99"`, sent as a string. No float rounding. | | enum | A string constrained to a declared set of variants. | | text | Unbounded UTF-8 prose. | | json | A structured JSON object or array, stored verbatim. | | list | A homogeneous array with an optional element type. | | inet | An IPv4 or IPv6 address. | | point | A geographic point - `{lat, lng}` or `[lng, lat]`. | ## Register a schema. Once you've written your `manifest.toml`, send it to OriginChainDB. After this succeeds, the table exists and you can start saving rows. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/schemas" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: text/plain" \ --data-binary @manifest.toml ``` #### Python ``` with open("manifest.toml") as f: db.schemas.register(f.read()) ``` #### TypeScript ``` import { readFileSync } from "node:fs"; await db.registerSchema(readFileSync("manifest.toml", "utf-8")); ``` #### Go ``` tomlBytes, _ := os.ReadFile("manifest.toml") req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/schemas", bytes.NewReader(tomlBytes)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "text/plain") http.DefaultClient.Do(req) ``` common mistakes - Wrong Content-Type. The body is TOML, not JSON. Use `Content-Type: text/plain`. - Forgetting primary_key. Every schema needs at least one column listed in the top-level `primary_key = ["..."]` array. - Wrong column type. Use the exact spellings above — the six core scalars and the purpose-built types are all real. What is not accepted is the familiar alias: `string`, `int` and `ulid` return 400 (write `str`, `i64` or `u64`, and `str` instead). `timestamp`, `decimal` and `vector` do work — see the [reference](https://originchaindb.com/docs/schemas/reference#columns) for the full set. ## What you can add beyond the basics. The minimal schema above lets you save and read rows. Add any of these blocks when you need more: [[[indexes]] Make filters fast Without an index, `WHERE status = ...` scans every row.](https://originchaindb.com/docs/schemas/reference#indexes) [[[relations]] Turn a column into a graph edge "Customers of a supplier", "friends of friends" - declare once, walk forever.](https://originchaindb.com/docs/schemas/reference#relations) [[[foreign_keys]] Refuse rows that point at nothing Reject an order with a non-existent customer at write time.](https://originchaindb.com/docs/schemas/reference#foreign-keys) [[[check_constraints]] Reject invalid rows `qty > 0`, `status IN ('paid','pending')`, etc.](https://originchaindb.com/docs/schemas/reference#check-constraints) [[[extractions]] Derive a column from nested JSON Pull `customer.address.country` into a queryable column at write time.](https://originchaindb.com/docs/schemas/reference#extractions) [migrations Add or rename columns later Online, no downtime. Bump version, write the new manifest, ship it.](https://originchaindb.com/docs/schemas/transactions) ## What is not in the schema. Two things people expect to declare on the schema but don't: - Vector embeddings. There is no `[[extractions.vector]]` block. The vector values are stored on a separate runtime endpoint (`POST /vector/:table/put`) and reference rows by primary key. What the schema can declare is the collection's shape — an optional [`[vector]` block](https://originchaindb.com/docs/schemas/reference#vector) pinning `dim` and `distance`. See [Vector tables](https://originchaindb.com/docs/schemas/vector). - Full-text indexes. There is no `[[extractions.fts]]` block. Full-text indexes live on their own runtime endpoint (`POST /fts/:table/:field`). Tokenizer and analyzer pipeline live on the index call, not the schema. See [FTS tables](https://originchaindb.com/docs/schemas/fts). The schema is purely the row contract. Vector and full-text are auxiliary indexes that link back to rows by primary key. ## Where to go next. [tutorial Build your first schema step by step A guided walk-through where you add each block one at a time, with the "why" alongside.](https://originchaindb.com/docs/schemas/tutorial) [reference Every field, every option The complete manual - indexes, relations, foreign keys, CHECK, extractions, migrations.](https://originchaindb.com/docs/schemas/reference) --- # Schema for Cypher - OriginChainDB Canonical source: https://originchaindb.com/docs/schemas/cypher Sitemap last modified: 2026-09-23T08:33:23.000Z schema · cypher # Cypher. Define the tables and relations needed to match graph patterns with supported Cypher syntax. Reach for it when the question is "what is this row connected to, and what is that connected to". A three-hop question is three arrows in Cypher and three joins in SQL, and the arrows stay readable. For aggregation, string matching, arithmetic or anything with a `GROUP BY` in it, use [SQL](https://originchaindb.com/docs/schemas/sql) - this dialect deliberately does not compete there. one route `POST /v1/tenants/:t/cypher` with a body of `{ cypher, default_schema?, params? }`. Reads and writes both go here. No SDK wraps this route yet, so every tab on this page builds the request by hand - which is three lines of boilerplate and then plain query strings. ## Before you start: the graph shape. A relationship is not a table. It is a column on the row that holds another row's primary key, plus a `[[relations]]` block naming it. Declare the block and every write to that table starts maintaining the edge for you. [Graph](https://originchaindb.com/docs/schemas/graph) walks through the modelling in depth; here is the shape these examples use. Register `shop.customers` first - a relation's target table must already exist when the pointing table is registered. It refers to itself through `referred_by`, which is what makes a variable-length referral chain legal later on. schemas/customers.toml ``` namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" [[columns]] name = "country" ty = "str" [[columns]] name = "referred_by" ty = "str" # another customer's id - or absent # A customer points at the customer who referred them. Source and target # are the SAME table, which is what makes variable-length paths legal. [[relations]] name = "referrer" from_col = "referred_by" target = { namespace = "shop", table = "customers", pk = "id" } bidirectional = true # REQUIRED for MATCH (c:customers {id: '...'}) to resolve. [[indexes]] name = "by_id" columns = ["id"] ``` schemas/orders.toml ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" # a shop.customers id [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # The edge: an order points at the customer who placed it. [[relations]] name = "placed_by" from_col = "customer" target = { namespace = "shop", table = "customers", pk = "id" } bidirectional = true [[indexes]] name = "by_id" columns = ["id"] ``` the by_id index is not optional `MATCH (o:orders {id: "..."})` resolves through an index named `by_`. Without that `[[indexes]]` block the point match does not resolve, and since every traversal has to start from a pinned node, nothing on this page will work. Declare it on any table you intend to query with Cypher. ## Related. [graph Graph modelling and algorithms How relations are declared and stored, plus the 22 graph endpoints.](https://originchaindb.com/docs/graph) [sql Schema for SQL Where to go for GROUP BY, string matching, arithmetic and joins.](https://originchaindb.com/docs/schemas/sql) [dashboard Run Cypher from the console The query workbench, with the same syntax and no HTTP boilerplate.](https://originchaindb.com/docs/dashboard/queries) [api HTTP API reference Auth, error envelopes, and every other route.](https://originchaindb.com/docs/api) --- # Full-text search - OriginChainDB schema reference Canonical source: https://originchaindb.com/docs/schemas/fts Sitemap last modified: 2026-09-23T08:33:23.000Z schema · full-text # Full-text search. Configure the full-text fields and analysis settings used to index and search document text. Reach for it when the user typed words and you want the words to matter: product search, log and ticket search, or the keyword half of a hybrid retrieval stack. When meaning matters more than wording, use [vector search](https://originchaindb.com/docs/schemas/vector); when you know the predicate exactly, use [SQL](https://originchaindb.com/docs/schemas/sql). Every operation below is shown in cURL and Python, and in TypeScript and Go where those SDKs wrap it. Where an SDK does not wrap something the tab says so and shows the raw call. 1 ## Before you start. Every example uses the `shop.orders` table from the [quickstart](https://originchaindb.com/docs/quickstart), full-text indexing its `notes` column. indexing is explicit — inserting a row does not index it This is the single most common surprise on this page. Writing a row through the rows endpoint or SQL does not put anything into the full-text index. You POST each document to the FTS endpoint yourself, and you do it again whenever the text changes. Nothing in the schema TOML marks a column as searchable, because the index is keyed on a free-form `(table, field)` pair rather than on your schema at all. The `:table` and `:field` path segments are opaque strings. They must match exactly between the write and the read — index into `shop.orders/notes` and search `shop.orders/note` and you get an empty result, not an error. Registering the row schema anyway is what lets you turn a `doc_id` back into a row. schemas/orders.toml ``` # Nothing in the schema marks a field as full-text indexed - there is no # such flag. An FTS index comes into existence the first time you POST a # document to a (table, field) pair. Registering the row schema is still # worth it: it is what lets you take a doc_id from a search hit and read # the whole row back with SQL. namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[columns]] name = "notes" # the field we will full-text index ty = "str" [[columns]] name = "placed_ms" ty = "u64" ``` 2 ## Index a document. One document, one call. Use the row's primary key as the `doc_id` so a hit maps straight back to a row. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "text": "rush delivery, signed by recipient" }' # → 201 Created, empty body ``` #### Python ``` db.fts.index( "shop.orders", "notes", doc_id="01JTRX9KQ3YH8K2WMX0F5JZAB7", text="rush delivery, signed by recipient", ) ``` #### TypeScript ``` await db.ftsIndex("shop.orders", "notes", { doc_id: "01JTRX9KQ3YH8K2WMX0F5JZAB7", text: "rush delivery, signed by recipient", }); ``` #### Go ``` err := db.FTSIndex(ctx, "shop.orders", "notes", originchain.FTSIndexRequest{ DocID: "01JTRX9KQ3YH8K2WMX0F5JZAB7", Text: "rush delivery, signed by recipient", }) ``` - Returns `201 Created` with an empty body. - The write is synchronous and atomic — postings, document length, token set and corpus statistics all land in one batch. The document is searchable the moment the call returns; a crash mid-call leaves every record or none. - Re-indexing the same `doc_id` replaces the previous version cleanly. There is no separate "update" or "delete from index" call — write it again. indexing a nested json document If your text is spread across a nested object, the `/json` variant walks it for you. Dotted paths select what to index; string arrays under a listed path flatten one level. No SDK wraps this variant. ``` # Walk a nested document and index its string leaves. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes/json" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "doc_id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "json": { "note": "rush delivery", "shipping": { "instructions": "signed by recipient" }, "tags": ["priority", "insured"] }, "paths": ["note", "shipping.instructions", "tags"] }' # Omit "paths" to index every string leaf in the document. ``` 3 ## Analysis, synonyms and stopwords. The analyzer is fixed: Unicode word segmentation, then lowercase. That is the whole pipeline, and there is no parameter to change it. no stemming today Searching `deliver` will not match `delivery` or `delivered` — they are three distinct terms. A stemmer covering eighteen languages exists inside the engine, but it is not selectable through the API, so today's behaviour is exact-token matching. Work around it with [fuzzy matching](https://originchaindb.com/docs/fts#fuzzy), a synonym class, or a `prefix` clause in the DSL. What you can configure, per `(table, field)` pair, is a synonym map and a stopword list. Both apply at index and query time, so installing either after you have indexed documents means re-indexing them to get consistent behaviour. #### cURL ``` # Synonym classes - applied at BOTH index and query time. # Re-installing replaces the whole map. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes/synonyms" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "synonyms": { "delivery": ["shipment", "dispatch"] } }' # Stopwords - dropped at BOTH index and query time. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/fts/shop.orders/notes/stopwords" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "stopwords": ["the", "a", "an", "and", "by", "of", "with"] }' ``` #### Python ``` db.fts.install_synonyms( "shop.orders", "notes", {"delivery": ["shipment", "dispatch"]}, ) db.fts.install_stopwords( "shop.orders", "notes", ["the", "a", "an", "and", "by", "of", "with"], ) ``` Each install replaces the whole map or list — there is no incremental add. A term may have at most 32 synonyms. Only the Python SDK wraps these two calls. ## Related. - [Full-text search — the modality reference](https://originchaindb.com/docs/fts) — search modes, ranking, highlights and the JSON query DSL. - [Quickstart — your first search](https://originchaindb.com/docs/quickstart#fts) in all four languages. - [Full-text examples](https://originchaindb.com/docs/examples/fts) — one focused page per query shape. - [Vector search](https://originchaindb.com/docs/schemas/vector) — the semantic half of a hybrid retrieval stack. - [Search from the dashboard](https://originchaindb.com/docs/dashboard/queries) — the console's one-line shorthand. - [Full schema reference](https://originchaindb.com/docs/schemas/reference) — every block the TOML grammar accepts. --- # Graph relations - OriginChainDB schema reference Canonical source: https://originchaindb.com/docs/schemas/graph Sitemap last modified: 2026-09-23T08:33:23.000Z schema · graph # Graph. Declare relationships between tables so existing rows can be traversed as a graph. Use it when the interesting part of a question is the connection rather than the value - fraud rings, referral trees, recommendation neighbourhoods, dependency chains, "who else touched this". If your question is really an aggregate with a join in it, [SQL](https://originchaindb.com/docs/schemas/sql) will be both faster and clearer. ## The one modelling decision that matters. This is where almost everyone goes wrong on their first schema, so it is worth being blunt about: An edge is a column on the node table that holds the other row's primary key. It is not a separate edge table with source and destination columns. If you have used a property-graph database, the instinct is to build a join table. Here that produces a table nothing can traverse: ``` # The instinct from other graph databases - DON'T do this. namespace = "shop" table = "order_customer_edges" primary_key = ["id"] [[columns]] name = "id" ty = "str" [[columns]] name = "src" # order id ty = "str" [[columns]] name = "dst" # customer id ty = "str" # This is just a table. Nothing traverses it. Every graph endpoint # and every Cypher arrow will refuse to touch it, because no # [[relations]] block points anywhere. ``` The working version puts the pointer on `shop.orders` itself. Register the target table first - a relation whose target does not exist yet fails validation. schemas/customers.toml — register this first ``` namespace = "shop" table = "customers" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" [[columns]] name = "country" ty = "str" [[columns]] name = "referred_by" ty = "str" # another customer's id # A self-relation: customers point at customers. [[relations]] name = "referrer" from_col = "referred_by" target = { namespace = "shop", table = "customers", pk = "id" } bidirectional = true [[indexes]] name = "by_id" columns = ["id"] ``` schemas/orders.toml ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" # <- THIS column is the edge ty = "str" [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # The declaration that turns a plain column into a traversable edge. [[relations]] name = "placed_by" # the name you pass as ?rel= from_col = "customer" # the column ON THIS TABLE target = { namespace = "shop", table = "customers", pk = "id" } bidirectional = true # default; enables reverse traversal [[indexes]] name = "by_id" columns = ["id"] ``` | Field | What it means | | --- | --- | | name | The edge's verb. This is the string you pass as `?rel=` on every graph endpoint and inside every Cypher arrow. It is not the column name. | | from_col | The column on this table whose value identifies the far row. The single most misread field on the page - it is a local column, not the target's. | | target | `{ namespace, table, pk }`. `pk` must be the target's single primary-key column. Cross-namespace targets are fine. | | bidirectional | Defaults to `true`. Writes a reverse edge as well, which is what makes `/reverse` and backward Cypher arrows work. Set it to `false` and those return empty rather than erroring. | `` ``There is no edge API Once the relation is declared, you never write an edge. You write an ordinary row, and the engine derives the forward and reverse edge keys inside the same commit. Delete the row and they go with it; change the column and the edge moves. cURL TypeScript Python Go POST /v1/tenants/:t/rows/shop.orders # Write an ORDINARY row. There is no edge API and nothing else to call. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "amount_cents": 12950, "status": "paid", "notes": "rush delivery", "placed_ms": 1714478049000 }' # The edge order -> customer now exists. Both directions, because # placed_by declares bidirectional = true.// The TS SDK doesn't wrap row writes yet - plain fetch. await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/shop.orders`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ id: "01JTRX9KQ3YH8K2WMX0F5JZAB7", customer: "01JTRX1H4Q9P0N2WMX0F5JZ001", amount_cents: 12950, status: "paid", notes: "rush delivery", placed_ms: 1714478049000, }), }, ); // The edge exists now. No separate edge write.db.rows.put("shop.orders", { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "amount_cents": 12950, "status": "paid", "notes": "rush delivery", "placed_ms": 1714478049000, }) # The edge exists now. No separate edge write.// The Go SDK doesn't wrap row writes yet - net/http. body, _ := json.Marshal(map[string]any{ "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "amount_cents": 12950, "status": "paid", "notes": "rush delivery", "placed_ms": uint64(1714478049000), }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/rows/shop.orders", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") http.DefaultClient.Do(req) // The edge exists now. No separate edge write. many-to-many from one column If from_col holds a JSON array instead of a single value, the engine emits one edge per element. That is how you model a genuine many-to-many - an order with several tags, a document with several authors - without ever creating a join table. A missing or null value emits no edge at all. One more thing worth knowing: the FK column does not need a secondary index. Edges live in their own key space, derived at write time, and traversal is a prefix scan over that space rather than a lookup on the column. The by_id index in the TOML above is there for Cypher, which anchors patterns by primary key - the graph endpoints do not need it.`` `Limits and gotchas. Primary keys must be single-column strings. The graph endpoints encode pk, src and dst as strings, so a table with an integer or composite primary key cannot be addressed through them at all. Model graph node ids as str from the start - ULIDs are the usual choice. The REST endpoints are single-schema. Each call resolves against one schema's catalog, so a traversal that has to cross into another namespace is not expressible here. Declaring a cross-namespace relation is fine - traversing one needs Cypher or a plan query. Turning bidirectional off is a one-way door in practice. The reverse edges are written at row-write time, so flipping the flag later does not backfill them for rows already stored. Rewrite the rows if you change your mind. Related. cypher Cypher Pattern syntax over the same relations, returning full rows. reference Graph endpoint reference Every parameter and tuning knob, endpoint by endpoint. schemas Schema reference Every TOML block, including relations, in one place. dashboard Design a schema visually Draw relations between tables on a canvas instead of writing TOML. ← Cypher Next: Schema for Vector →` --- # Materialized views and refresh modes - OriginChainDB Canonical source: https://originchaindb.com/docs/schemas/materialized-views Sitemap last modified: 2026-09-13T20:18:50.000Z schema · materialized views # Materialized views. A materialized view is a SELECT whose result is computed once and stored. Reads hit the stored rows instead of re-scanning the base table, so a dashboard aggregate that costs a full scan every time it loads costs one key lookup instead. Views are a runtime object, not a schema object. Nothing about them lives in your table's TOML - you install one over an existing table through three HTTP routes, and the base table is untouched. the whole surface - `POST /v1/tenants/:t/sql/materialized-views` - install - `GET /v1/tenants/:t/sql/materialized-views/:name` - read the rows - `POST /v1/tenants/:t/sql/materialized-views/:name/refresh` - recompute That is the complete list. There is no list route, no drop route, and no `CREATE MATERIALIZED VIEW` in SQL - the SQL translator rejects it explicitly. ## When a view beats a plain query. A view is a cache with a manual invalidation button. It pays off when the read/write ratio is lopsided and a little staleness is acceptable. | Use a view when | Use a plain query when | | --- | --- | | The same aggregate is read many times between writes - a dashboard tile, a leaderboard, a nightly rollup. | Every read has different parameters. A view stores one fixed result set, not a parameterised one. | | The base table is large enough that the scan dominates response time. | The answer must be exactly current on every read and you are not refreshing on every write. | | You control when the data changes, so you know exactly when to refresh (end of an import, end of a billing period). | Writes are constant and reads are rare - you would spend more on refreshes than you save on reads. | ## Before you start. Every example below runs against the `shop.orders` table from the [quickstart](https://originchaindb.com/docs/quickstart). Register it first if you haven't - views are installed over schemas that already exist, and nothing in this TOML is view-specific. schemas/orders.toml ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" # money in minor units - never f64 [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # epoch milliseconds # Makes the WHERE status = 'paid' inside the view cheap to re-run. [[indexes]] name = "by_status" columns = ["status"] ``` The query we will materialize is the "revenue per customer" aggregate - the kind of thing a dashboard asks for on every page load: ``` SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total_cents FROM shop.orders WHERE status = 'paid' GROUP BY customer ``` test the query first Install runs the query through the same translator as `POST /sql`. If the SELECT doesn't compile there, install returns the identical `400`. Get the query green on [/sql](https://originchaindb.com/docs/schemas/sql) first, then wrap it in a view. ## 1. Install a view. Install does three things in one call: it translates the SQL, runs the query once to produce the initial snapshot, and commits that snapshot in a single write-ahead-log frame. When the call returns `200`, the view is already populated - there is no separate backfill step and no window where the view exists but is empty. The body is `{ name, query, refresh_mode?, source_schema? }`. Omitting `refresh_mode` gives you `on_demand`; `source_schema` is a hint the engine otherwise derives from the plan's first scan target. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql/materialized-views" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "mv_paid_by_customer", "query": "SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total_cents FROM shop.orders WHERE status = '\''paid'\'' GROUP BY customer", "refresh_mode": "on_demand" }' ``` #### TypeScript ``` // No SDK wrapper for materialized views yet - plain fetch. const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/sql/materialized-views`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ name: "mv_paid_by_customer", query: "SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total_cents " + "FROM shop.orders WHERE status = 'paid' GROUP BY customer", refresh_mode: "on_demand", }), }, ); const out = await res.json(); console.log(out.rows_materialized, "groups materialized"); ``` #### Python ``` # The SDK helper (db.sql.install_materialized_view) still sends the # pre-release refresh_mode names, so call the endpoint directly. import os, requests BASE = f"https://{os.environ['OC_HOST']}/v1/tenants/{os.environ['OC_TENANT']}" H = {"Authorization": f"Bearer {os.environ['OC_TOKEN']}"} r = requests.post( f"{BASE}/sql/materialized-views", headers=H, json={ "name": "mv_paid_by_customer", "query": ( "SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total_cents " "FROM shop.orders WHERE status = 'paid' GROUP BY customer" ), "refresh_mode": "on_demand", }, ) r.raise_for_status() print(r.json()["rows_materialized"], "groups materialized") ``` #### Go ``` // No SDK wrapper for materialized views yet - net/http. body, _ := json.Marshal(map[string]any{ "name": "mv_paid_by_customer", "query": "SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total_cents " + "FROM shop.orders WHERE status = 'paid' GROUP BY customer", "refresh_mode": "on_demand", }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/sql/materialized-views", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var out struct { Name string `json:"name"` RowsMaterialized int `json:"rows_materialized"` BytesWritten int `json:"bytes_written"` RefreshTS uint64 `json:"refresh_ts"` } json.NewDecoder(resp.Body).Decode(&out) fmt.Println(out.RowsMaterialized, "groups materialized") ``` response ``` { "name": "mv_paid_by_customer", "rows_materialized": 3, "bytes_written": 412, "refresh_ts": 1714478100000 } ``` `rows_materialized` is the number of result rows stored - here, one per customer group, not the 12,480 base rows that were scanned to produce them. `refresh_ts` is the millisecond timestamp of the snapshot; it is the only "how stale is this" signal the engine gives you, so log it. ## 2. Read the view. A view is read by name through its own route - it is not addressable from SQL. You cannot write `SELECT * FROM mv_paid_by_customer`; the view is not a table and the planner does not know its name. Read it, then use the rows in your application. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/sql/materialized-views/mv_paid_by_customer" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/sql/materialized-views/mv_paid_by_customer`, { headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}` } }, ); const view = await res.json(); for (const row of view.rows) { console.log(row.customer, row.orders, row.total_cents); } ``` #### Python ``` view = db.sql.read_materialized_view("mv_paid_by_customer") for row in view.rows: print(row["customer"], row["orders"], row["total_cents"]) ``` #### Go ``` req, _ := http.NewRequestWithContext(ctx, "GET", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/sql/materialized-views/mv_paid_by_customer", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var view struct { Name string `json:"name"` Rows []map[string]any `json:"rows"` } json.NewDecoder(resp.Body).Decode(&view) for _, row := range view.Rows { fmt.Println(row["customer"], row["orders"], row["total_cents"]) } ``` response ``` { "name": "mv_paid_by_customer", "rows": [ { "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001", "orders": 4, "total_cents": 51800 }, { "customer": "01JTRX1H4Q9P0N2WMX0F5JZ002", "orders": 2, "total_cents": 18400 }, { "customer": "01JTRX1H4Q9P0N2WMX0F5JZ003", "orders": 1, "total_cents": 9900 } ] } ``` The row shape is exactly what the original SELECT projected, aliases included - `orders` and `total_cents` here come straight from the `AS` clauses. The whole snapshot comes back in one response; there is no pagination on this route, which is the practical ceiling on how many groups a view should have. ## 3. Refresh the view. Under `on_demand` - the default and the mode to build against - the snapshot never changes on its own. Writes to `shop.orders` do not touch the view. It goes stale the instant the first write lands, and it stays exactly as stale as it was until you call refresh. Refresh reloads the stored definition, re-translates the SQL against the current catalog, re-executes it, and atomically overwrites the snapshot. It is a full recompute, not a delta: cost scales with the base table, not with how much changed. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql/materialized-views/mv_paid_by_customer/refresh" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### TypeScript ``` const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/sql/materialized-views/mv_paid_by_customer/refresh`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}` }, }, ); const out = await res.json(); console.log(out.rows_materialized, "groups rebuilt at", out.refresh_ts); ``` #### Python ``` out = db.sql.refresh_materialized_view("mv_paid_by_customer") print(out.rows_materialized, "groups rebuilt at", out.refresh_ts) ``` #### Go ``` req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/sql/materialized-views/mv_paid_by_customer/refresh", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var out struct { Name string `json:"name"` RowsMaterialized int `json:"rows_materialized"` RefreshTS uint64 `json:"refresh_ts"` } json.NewDecoder(resp.Body).Decode(&out) fmt.Println(out.RowsMaterialized, "groups rebuilt at", out.RefreshTS) ``` response ``` { "name": "mv_paid_by_customer", "rows_materialized": 4, "bytes_written": 540, "refresh_ts": 1714481700000 } ``` nothing schedules this for you There is no refresh interval, no cron, no auto-refresh setting. If your view should be at most five minutes stale, something on your side has to call this route every five minutes. The usual pattern is to refresh at the end of the job that writes the data, so the view is fresh precisely when it matters. ## Refresh modes, precisely. `refresh_mode` takes exactly two values. Anything else is a hard error, not a fallback - a typo does not quietly become the default. | Mode | Status | What it does | | --- | --- | --- | | on_demand | Default. Shipped. | The snapshot changes only when you POST to `/refresh`. This is the one mode you can plan a product around. | | incremental | Preview, gated per kind. | The view is maintained apply-time - updated as each base write commits, with no refresh call. Availability depends on server-side gates; see below. | ### What "apply-time" actually means When an incremental view is live, maintenance is not a background job and not a lagging follower. The engine updates the view's stored cells while holding the same store write lock the base commit takes, so the base row and the view move together: once a write returns, a read of the view already reflects it. There is no replication delay to reason about and no window where the two disagree. Install closes the obvious race too. The initial fill and the switch that turns on live maintenance happen under one lock hold, so a write that lands in between cannot be counted twice or dropped - which would otherwise be a permanent miscount that only a full refresh could repair. The trade is where the cost sits. `on_demand` makes writes free and pays the whole scan on refresh. `incremental` puts a small amount of work on the commit path of every write to the base table, and reads are always current. A write-heavy table with a rarely-read view is the wrong place for it. ### Why an incremental install is usually refused Incremental maintenance ships behind two independent preview gates - one for filter/projection views, one for aggregate views - and an install is only accepted if its own kind's gate is on. The filter/projection gate is on by default; the aggregate gate is off by default. So the revenue-per-customer view above is refused rather than silently maintained wrong: #### Request ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql/materialized-views" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "mv_paid_live", "query": "SELECT customer, SUM(amount_cents) AS total_cents FROM shop.orders WHERE status = '\''paid'\'' GROUP BY customer", "refresh_mode": "incremental" }' ``` #### 422 ``` HTTP/1.1 422 Unprocessable Entity { "error": "incremental materialized views over a COUNT/SUM/AVG aggregate are a preview feature that is DISABLED on this server (OC_MV_AGG_PREVIEW); use refresh_mode=on_demand, or ask your operator to enable the preview" } ``` #### 400 ``` HTTP/1.1 400 Bad Request { "error": "incremental materialized views (preview) support only a pure filter/projection over a single table - no MIN/MAX, DISTINCT, join, set-op, subquery, or limit; use refresh_mode=on_demand for this query" } ``` The distinction is worth reading closely. A 422 means "this shape is supported but switched off here" and names the gate. A 400 means "this shape is not supported by incremental at all". Both tell you to use `on_demand`, and both are honest refusals - the engine will not accept an incremental view it cannot keep correct. shapes incremental can maintain Only two, and only over a single table: - a pure filter / projection; - a `COUNT` / `SUM` / `AVG` aggregate, with an optional `GROUP BY`. Explicitly excluded: `MIN`, `MAX`, `DISTINCT`, joins, set operations, subqueries, `HAVING`, `ORDER BY` and `LIMIT`. `MIN` and `MAX` are not an oversight: retracting the current extreme would require remembering the next one, which an incremental counter does not keep. `on_demand` has none of these restrictions - it accepts any SELECT the translator accepts, joins and window functions included, because it simply re-runs the query. two traps that surprise people An index on the filtered column disqualifies the view. Incremental classification requires the view's plan to sit directly on a table scan. If `status` has an index, `WHERE status = 'paid'` plans as an index scan instead, and the same SQL that installed yesterday stops classifying. Whether an incremental install succeeds can therefore depend on your indexes, not just your SQL. A bare `SELECT COUNT(*) FROM shop.orders` is not eligible. With no `WHERE` and no `GROUP BY` it compiles to a dedicated count operator rather than a general aggregate, which the incremental classifier does not recognise. Add the `WHERE` you almost certainly wanted, or use `on_demand`. One safety property worth knowing: if an incremental view ever stops being maintained - the gate is turned off, or a rebuild fails after a restart - reads do not serve its stale cells. The engine recomputes from live base rows for that read instead, so a read is never wrong, only more expensive than you expected. A successful refresh, or a restart with the gate back on, restores fast serving. ## Limits and gotchas. - A view cannot be dropped. There is no DELETE route. Installing over an existing name returns `409`, so a name is claimed permanently and a view's query can never be edited. Version the name if you expect the definition to change - `mv_paid_by_customer_v2` - rather than hoping to replace it. - A view cannot be listed. No index route exists, and reads are by exact name. Keep view names in source control; a forgotten name is unreachable and unremovable. - Views are invisible to SQL. They cannot appear in a `FROM` clause, cannot be joined, and cannot be nested inside another view. `CREATE MATERIALIZED VIEW` is rejected by the translator by design - the HTTP route is the only way in. - The snapshot is returned whole, and it is capped at 16 MiB. There is no `LIMIT`, offset, or filter on the read route, and an `on_demand` snapshot that encodes larger than 16 MiB is rejected with a `400` at install and at refresh - so a view can outgrow its cap months after it was created. Group by something with bounded cardinality; customer is fine, a raw millisecond timestamp is not. Nothing caps group count directly, so an unbounded `GROUP BY` is your problem to avoid. - A view cannot read another view. Installing a view whose query references an existing view's name is rejected. Views do not compose. - Single-shard only. On a sharded instance, a view whose base tables do not live on the same shard as the view itself returns `501` at install and at refresh. Cross-shard views are not in this release. - Dropping the base table does not clean up an `on_demand` view. The snapshot survives and keeps serving its last contents, the name stays claimed, and with no drop route there is no way to remove either. Drop the view's data source only when you have accepted that. - Incremental `SUM` and `AVG` can drift in the low bits. Floating-point addition is not associative, so a value folded write-by-write can differ slightly from the same value computed by one scan. `COUNT` is exact. A refresh rebuilds from scratch and reconciles the difference - schedule one periodically if you display these figures to the decimal. - Refresh is a full recompute. Changing one row costs the same refresh as changing a million. Refreshing an `on_demand` view in a per-write hook is the anti-pattern this design invites - batch it. - Names are unique per view, not per query. Nothing stops two views from materializing the same SELECT under different names, and nothing deduplicates the work. Both will need refreshing. - The Python SDK helper is behind the engine. `db.sql.install_materialized_view()` still sends the pre-release mode names `manual` / `on_write`, which the engine no longer recognises. Call the install route directly, as the Python tab above does. `read_materialized_view()` and `refresh_materialized_view()` take no mode and work as documented. - TypeScript and Go have no view helpers. Both tabs above use the raw HTTP client. The routes are stable; the wrappers are not shipped. ### Status codes you will actually hit | Code | When | What to do | | --- | --- | --- | | 400 | The view's SQL doesn't translate - unknown column, unknown schema, or a shape the translator refuses. | Run the query against POST /sql first. If it 400s there, it 400s here. | | 400 | refresh_mode is a string the engine doesn't know (anything other than on_demand or incremental). | Unknown values are a hard error, never a silent fallback. Omit the field to get on_demand. | | 409 | A view with that name is already installed. | Pick a different name. Install is not an upsert and there is no drop route. | | 404 | Refreshing or reading a name that was never installed. | Check the name. There is no list endpoint, so keep your view names in source control. | | 422 | refresh_mode: "incremental" on a server where that kind's preview gate is off. | Use on_demand. The message names the gate the operator would have to turn on. | ## Related. [sql Schema for SQL Which SELECT shapes translate - test your view's query here first.](https://originchaindb.com/docs/schemas/sql) [explain Read a query plan Find out whether the scan really is the bottleneck before you cache it.](https://originchaindb.com/docs/dashboard/explain) [transactions Schema for transactions How writes commit - the same path an incremental view is maintained on.](https://originchaindb.com/docs/schemas/transactions) [api HTTP API reference Auth headers, error envelopes, and every other route.](https://originchaindb.com/docs/api) --- # Natural language schemas - OriginChainDB schema reference Canonical source: https://originchaindb.com/docs/schemas/nl Sitemap last modified: 2026-09-23T08:33:23.000Z schema · natural language # Natural language. Make your registered schemas available to natural-language queries and inspect the resulting plan. It earns its place in two situations: exploration, when you don't yet know the shape of the data well enough to write the query, and end-user surfaces, where the person asking will never write a query at all. For anything on a hot path or in a code path you will maintain, write the [SQL](https://originchaindb.com/docs/schemas/sql) yourself — see [when to use which](https://originchaindb.com/docs/nql/quickstart#when). 1 ## Before you start. Every example uses the `shop.orders` table from the [quickstart](https://originchaindb.com/docs/quickstart). NL adds no schema syntax of its own — what it needs is that your tables are registered, because the column declarations are the entire context the compiler gets. schemas/orders.toml ``` # NL has no schema knobs of its own. What it needs is that the tables you # want it to reason about are REGISTERED - the compiler is given the column # names and types of your schemas as its entire context, so a column that # is not declared is a column it cannot use. namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # Not required by NL, but the compiled plan will use it like any other # query would - so an indexed filter column makes the answer faster. [[indexes]] name = "by_status" columns = ["status"] ``` column names are your prompt The compiler sees column names and types — nothing else. No sample rows, no comments, no descriptions. That makes naming load-bearing: `amount_cents` tells it far more than `amt`, and `placed_ms` tells it more than `ts`. If a question keeps compiling wrong, look at your column names before you look at your phrasing. ## Related. - [The ask endpoint](https://originchaindb.com/docs/ask) — the full HTTP reference. - [Natural-language examples](https://originchaindb.com/docs/examples/ask) — one focused page per scenario. - [SQL](https://originchaindb.com/docs/schemas/sql) — what to write once the question stops changing. - [Asking from the dashboard](https://originchaindb.com/docs/dashboard/queries) — NL in the query workbench. - [Full schema reference](https://originchaindb.com/docs/schemas/reference) — every block the TOML grammar accepts. --- # Schema reference - every field, every option Canonical source: https://originchaindb.com/docs/schemas/reference Sitemap last modified: 2026-09-13T20:18:50.000Z reference · schemas # Schema reference An OriginChainDB schema is a TOML manifest with three required top-level keys - namespace, table and primary_key. Columns, indexes, relations, foreign keys, CHECK constraints, JSON extractions and row expiry are declared as blocks beneath them. Use the [tutorial](https://originchaindb.com/docs/schemas/tutorial) if you want to learn it for the first time. ## 1. Top-level fields. | Field | Type | Required | Notes | | --- | --- | --- | --- | | namespace | string | yes | Letter-prefixed, snake_case. Groups related tables. | | table | string | yes | Letter-prefixed, snake_case. The table's name. Queries address it as `namespace.table`. | | primary_key | string[] | yes | One or more column names that uniquely identify a row. Composite keys are supported in SQL queries; the `/rows/:schema/:pk` shortcut is single-column only. | | version | int | no | Monotonic counter. Defaults to `1`. Schema migrations require an incremented value - see [migrations](https://originchaindb.com/docs/schemas). | ## 2. Columns. One `[[columns]]` block per column. Three fields per block. | Field | Type | Notes | | --- | --- | --- | | name | string | The column name. Snake_case. Used as the JSON key in row writes. | | ty | enum | Core scalars: `str`, `i64`, `u64`, `f64`, `bool`, `bytes`. Extended types: `timestamp`, `date`, `uuid`, `decimal`, `enum`, `text`, `json`, `list`, `inet`, `interval`, `point`, `vector` — see the table below. | | required | bool | Default `false`. When `true`, writes that omit this column are rejected. | ### Extended column types Beyond the six core scalars, OriginChainDB accepts purpose-built types. Each is format-validated at write time — a value of the wrong shape is rejected with a `400` type-mismatch. They store and index on the same substrate as their backing scalar, so range scans and filters stay fast. | Type | Wire form | Notes | | --- | --- | --- | | timestamp | integer (epoch ms) | Milliseconds since 1970-01-01 UTC. Indexed and range-scanned like an integer — date ranges are fast. | | date | integer (epoch days) | Whole days since 1970-01-01, no time component. | | interval | integer (ms) | A duration in whole milliseconds; may be negative. | | uuid | string | Canonical 8-4-4-4-12 hex form; validated on write. | | decimal | string | Exact fixed-point (e.g. `"19.99"`) — no binary-float rounding, money-safe. Optional `precision` / `scale` are enforced at write. Sent as a string, not a number. | | enum | string | Constrained to a declared `variants` list; a value outside the set is rejected. | | text | string | Unbounded UTF-8 prose. Behaves like `str`. | | json | object / array | A structured JSON document, stored and returned verbatim. | | list | array | A homogeneous array; an optional `element` type validates every item. | | inet | string | An IPv4 or IPv6 address; parsed and validated on write. | | point | object / pair | A geographic point: `{"lat": …, "lng": …}` or `[lng, lat]`. | | vector | not in the row body | A dense embedding. The column must declare its width in `[columns.params]` as `vector_dim`; a vector column without it is rejected. The values themselves live in a separate keyspace and are written through the vector endpoints, so the row keeps only the id — see [vector collections](https://originchaindb.com/docs/schemas/reference#vector). | ### Picking the right type for common data | Your data | Use | Why | | --- | --- | --- | | User ID, order ID, ULID | str | Travels as canonical text form. For RFC-4122 UUIDs use `uuid` to get format validation. | | Email, name, category | str | Lowercase emails at write time if you want case-insensitive equality. | | Price, amount, balance | decimal | Exact fixed-point — no float rounding. Send as a string (`"19.99"`). Or keep using `i64` minor units (cents, paise) if you prefer integers. | | Quantity, refundable count | i64 | Signed so refunds can be negative. | | Event time, created/updated at | timestamp | Epoch milliseconds. Range-scans like an integer; the planner knows it's a time so date filters stay fast. | | Rate, percentage | f64 | Float is fine for ratios where rounding doesn't matter. | | Active flag, deleted flag | bool | Single bit. | | Compressed blob, encoded payload | bytes | Opaque to the engine. | ## 3. Indexes. One `[[indexes]]` block per index. Without indexes, every filter does a full table scan. ``` [[indexes]] name = "by_status" columns = ["status"] # single-column [[indexes]] name = "by_customer_and_time" columns = ["customer_id", "placed_ms"] # composite, in declared order ``` | Field | Notes | | --- | --- | | name | A name for the index. Used in EXPLAIN output. | | columns | List of column names. Composite indexes serve left-prefix matches - `[customer_id, placed_ms]` can serve `WHERE customer_id = ?` alone, but not `WHERE placed_ms = ?` alone. | When to add an index: any column you filter on in `WHERE` or use as a foreign-key target. Indexes cost write throughput - don't add them speculatively. ## 4. Relations (graph edges). A relation turns a foreign-key column into a graph edge. The forward and reverse edges are written atomically when you save a row. ``` [[relations]] name = "supplied_by" # the verb you'll walk in code from_col = "supplier_id" # the FK column on THIS table bidirectional = true # default; reverse edge written atomically [relations.target] namespace = "shop" # target table's namespace table = "suppliers" # target table's name pk = "id" # target table's primary-key column # Self-relations work - direction tags resolve the collision: [[relations]] name = "follows" from_col = "followee_id" bidirectional = true [relations.target] namespace = "social" table = "users" pk = "id" ``` | Field | Notes | | --- | --- | | name | The edge verb - used as the `rel=` param when walking the edge with the graph endpoint. | | from_col | The column on this table that holds the target's primary key. | | bidirectional | Default `true`. When true, the reverse edge is also indexed - "who points at this row?" becomes a one-call lookup. | | [relations.target] namespace | Target table's namespace. | | [relations.target] table | Target table's name. | | [relations.target] pk | Target table's primary-key column. | Relations alone do not enforce that the target row exists. For that, also declare a foreign key (next section). ## 5. Foreign keys. Reject writes that point at non-existent rows. The integrity check is a single key fetch, so it's fast. ``` [[foreign_keys]] name = "fk_customer" # optional - defaults to "_fk" from_col = "customer_id" # column on THIS table target_schema = "shop.customers" # "namespace.table" of the target target_col = "id" # PK or unique-indexed column on target on_delete = "no_action" # no_action | restrict | set_null ``` | Field | Notes | | --- | --- | | name | Optional. Defaults to `"_fk"`. Used in error messages. | | from_col | Column on this table. | | target_schema | The target schema as `"namespace.table"`. | | target_col | Must be the target's primary key, or a unique-indexed column. | | on_delete | `no_action` (default), `restrict`, or `set_null`. `cascade` and `set_default` are not yet supported. | ## 6. CHECK constraints. Reject writes whose values violate a rule. Checks run on every write - keep them simple. ``` [[check_constraints]] name = "qty_positive" expression = "qty > 0" [[check_constraints]] name = "country_or_continent_set" expression = "country IS NOT NULL OR continent IS NOT NULL" [[check_constraints]] name = "status_in_set" expression = "status IN ('pending', 'paid', 'cancelled')" ``` ### Expression grammar | Operator | Example | | --- | --- | | = != < <= > >= | qty > 0 | | AND OR NOT | price > 0 AND price < 1000000 | | IS NULL / IS NOT NULL | description IS NOT NULL | | IN (...) | status IN ('pending', 'paid') | Three-valued logic applies: NULL on any operand yields NULL, which fails the CHECK. If a column can be NULL and you want the CHECK to ignore null values, use `col IS NULL OR col > 0`. ## 7. JSON extractions. Pull a value out of a nested JSON row into a queryable column at write time. Useful when your row body is already-structured JSON and you don't want to flatten it manually. ``` # Pull a derived column out of a nested JSON row at write time. # The dotted path walks into the JSON; missing fields emit zero values. [[extractions]] name = "customer_country" path = "customer.address.country" ty = "str" [[extractions]] name = "order_total_cents" path = "totals.grand_cents" ty = "i64" ``` | Field | Notes | | --- | --- | | name | The derived column's name. Queryable with SQL after extraction. | | path | Dotted JSON path walked at write time. Missing fields emit the type's zero value. | | ty | Any column type, core or extended. | Extractions are not the same as vector or full-text indexes. There is no `[[extractions.fts]]` or `[[extractions.vector]]` block - those live on separate runtime endpoints. ## 8. Row expiry (TTL). Give a table a time-to-live and its rows expire on their own. Once a row passes its expiry it is automatically excluded from every read - SELECTs, primary-key lookups, counts, vector and full-text search all skip it, as if it were gone - with no query changes and no clean-up job on your side. ``` # Row expiry (TTL) - two ways to declare it. Use ONE. # (A) Fixed lifetime: put `ttl` on any column and the WHOLE row expires # this long after it is written. Units: s / m / h / d / w. [[columns]] name = "id" ty = "str" required = true ttl = "30d" # every row lives 30 days from write # (B) Per-row expiry: instead of (A), add a [ttl] block pointing at a # timestamp or date column. Each row expires at its own stored value. # # [ttl] # column = "expires_at" # a `timestamp` or `date` column ``` ### Two ways to declare it | Declare | Meaning | | --- | --- | | ttl = "30d" | Fixed lifetime. Put it on any column; the whole row expires this long after it is written. Units: `s` `m` `h` `d` `w`. | | [ttl] column = "expires_at" | Per-row expiry. Point at a `timestamp` or `date` column; each row expires at its own stored value. Change the column and the expiry moves with it. | A row with no expiry set - no `ttl` column, or a null value in per-row mode - never expires. Rows written before you added a TTL keep living until you rewrite them. The row is the unit of expiry: a `ttl` written on one column still expires the whole row - every field, including its primary key and any foreign-key columns - not just that column. There is no way to expire a single column on its own; the primary key and referencing columns always leave together with the row, so a row never ends up half-gone. Expiry is per-table and does not cascade. If a row here is referenced from another table, letting it expire does not delete those referencing rows, and does not touch the row it points to - give related tables lifetimes that make sense together. (Referential actions run on the normal delete path, so when a table enforces foreign keys the expiry honours whatever `ON DELETE` rule you defined.) Over SQL, declare the same thing on `CREATE TABLE` - `WITH (ttl = '30d')` for a fixed lifetime, or `WITH (ttl_column = 'expires_at')` for per-row expiry: ``` CREATE TABLE app.sessions ( id TEXT NOT NULL PRIMARY KEY, user_id TEXT ) WITH (ttl = '30d'); ``` ## 9. Vector collections. One optional `[vector]` table per manifest. It is a single table, not a repeated `[[...]]` block, and it states what this table's embedding collection is: how wide its vectors are, which distance function they are indexed under, and optionally which model produced them. The engine then refuses a write or a query that disagrees, rather than returning a confident wrong ranking. Vector values are never stored in the row body. They live in their own keyspace keyed by `(tenant, table, id)` and are written through the [vector endpoints](https://originchaindb.com/docs/schemas/vector); the row keeps only the id. The `[vector]` block is metadata about that collection, not storage. ``` # Optional. One [vector] table per manifest. [vector] dim = 768 # REQUIRED. No default. distance = "cosine" # the default. Five values accepted. embedder = "BAAI/bge-large-en-v1.5" # optional model identity # A `vector` COLUMN carries the same width in its own params block. # Declare both and the two dims must agree, or registration fails. [[columns]] name = "embedding" ty = "vector" [columns.params] vector_dim = 768 ``` | Field | Type | Required | Notes | | --- | --- | --- | --- | | dim | int | yes | The dimensionality every stored and queried vector must match. There is no default — a `[vector]` block without it is rejected, as is a value of `0`. Enforced on every put and every top-k, from the moment you register it and before the first vector exists. | | distance | string | no | Defaults to `"cosine"`. Five values are accepted: `cosine`, `l2`, `dot`, `manhattan` and `l1` (an alias of `manhattan`). Anything else is refused at registration. See [how it resolves at runtime](https://originchaindb.com/docs/schemas/reference#vector-distance) below. | | embedder | string | no | The model identity the collection is indexed with, e.g. `"BAAI/bge-large-en-v1.5"`. When set, a put or a top-k that declares a different embedder is refused with a `400`; a caller that declares none is not blocked. When unset, only the dimension is enforced. Two models can agree on width and disagree on everything else, which is what this field is for. | ### How distance resolves at runtime `distance` is not documentation — it resolves and enforces the metric a write and a query actually run under. An index's neighbour lists are chosen by the distance function, so building under one metric and querying under another returns a plausible, wrongly-ordered answer that no health check can catch. | Declared | Request sends `metric` | Result | | --- | --- | --- | | Non-default (`l2`, `dot`, `manhattan`, `l1`) | a different one | 400 on both put and top-k. Not silently overridden in either direction. | | Any value | nothing | The declared metric is used — not `cosine`. Say nothing and you get what the collection says it is. | | `"cosine"` | something else | Allowed. `cosine` is the value an absent declaration also produces, so it is not evidence that anyone declared anything and cannot be enforced against. | | No `[vector]` block | anything or nothing | The request wins, defaulting to `cosine` when absent. Nothing is enforced. | The practical consequence: if your collection is anything other than cosine, declare it. A declared `"cosine"` is indistinguishable from no declaration at all, so it is the one value that documents intent without enforcing it. spell it exactly Schema registration does not reject unknown fields. A manifest carrying a block or key the engine has never heard of still registers with a `200` — it is simply ignored, every guard above stays off, and the response gives you no signal. So a typo here does not fail loudly; it fails silently, months later, as a wrong answer. - The table is `[vector]`, singular. There is no `[[vectors]]` array-of-tables, and the field is `distance`, not `metric` — `metric` is the name of the request field. - `index`, `nlist` and `quantization` are not manifest fields. Index kind, IVF centroids and quantization are chosen with install-time calls — see [IVF](https://originchaindb.com/docs/vector/ivf) and [quantization](https://originchaindb.com/docs/vector/quantization). - A vector column puts its width in `[columns.params]` as `vector_dim`. A bare `dim` on the column is one of the ignored-unknown-field cases. ## 10. Full example. Every block from this page in one schema. ``` # A schema using every block at once. namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer_id" ty = "str" required = true [[columns]] name = "product_id" ty = "str" required = true [[columns]] name = "qty" ty = "i64" [[columns]] name = "total_cents" ty = "i64" [[columns]] name = "placed_ms" ty = "u64" [[indexes]] name = "by_customer" columns = ["customer_id"] [[indexes]] name = "by_product_and_time" columns = ["product_id", "placed_ms"] [[relations]] name = "bought_product" from_col = "product_id" bidirectional = true [relations.target] namespace = "shop" table = "products" pk = "id" [[relations]] name = "by_customer" from_col = "customer_id" bidirectional = true [relations.target] namespace = "shop" table = "customers" pk = "id" [[foreign_keys]] name = "fk_customer" from_col = "customer_id" target_schema = "shop.customers" target_col = "id" on_delete = "restrict" [[check_constraints]] name = "qty_positive" expression = "qty > 0" [[check_constraints]] name = "total_positive" expression = "total_cents > 0" ``` --- # Row CRUD and batch ingest - OriginChainDB Canonical source: https://originchaindb.com/docs/schemas/row-crud Sitemap last modified: 2026-09-13T20:18:50.000Z query shapes · row crud # Row CRUD. The row endpoints are the direct door to the store: address a row by its primary key and write or read it, with no query to plan. Reach for them when your application already knows the key — fetching a session, saving a record, loading an order by id — and when you are loading data in bulk, where the batch route is by a wide margin the fastest path into the database. Use [SQL](https://originchaindb.com/docs/schemas/sql) instead when the question is "which rows", when you need a projection or an aggregate, or when you want a partial update. The two surfaces write to the same store and enforce the same constraints — but they differ on one thing that matters, covered in [section 4](https://originchaindb.com/docs/schemas/row-crud#upsert). 1 ## Before you start. Every example on this page uses the `shop.orders` table from the [quickstart](https://originchaindb.com/docs/quickstart). Two things in this manifest matter for the row endpoints specifically: - A single-column primary key. The `/rows/:schema/:pk` point read addresses one path segment. A table with a composite primary key is readable only through SQL. - Declared columns. They drive type validation and the canonical on-disk form of each value. ``` namespace = "shop" table = "orders" primary_key = ["id"] # single column - required for the /:pk shortcut [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" # money in minor units - never f64 [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # epoch milliseconds # Makes WHERE status = '...' sub-linear for SQL reads. [[indexes]] name = "by_status" columns = ["status"] # Enforced on EVERY write path, including the batch hot path. [[check_constraints]] name = "amount_positive" expression = "amount_cents > 0" ``` The shell examples assume `$OC_HOST`, `$OC_TENANT` and `$OC_TOKEN` are set — see [Authentication](https://originchaindb.com/docs/auth). The SDK examples assume a client named `db`, built as in the [quickstart](https://originchaindb.com/docs/quickstart). 2 ## The whole surface. Three routes. That is the complete list — there is no `PUT`, no `PATCH`, no `DELETE`, and no list route on the row surface. | Method | Path | Body | Response | | --- | --- | --- | --- | | POST | /v1/tenants/:t/rows/:schema | One row object, flat. | 200 { "ok": true } | | GET | /v1/tenants/:t/rows/:schema/:pk | — | 200 the row object, bare · 404 if absent | | POST | /v1/tenants/:t/rows/:schema/_batch | A raw array of rows, or newline-delimited JSON. | 200 { "inserted": n } · 207 on a partial stream | Deleting a row and reading a set of rows both go through [SQL](https://originchaindb.com/docs/schemas/sql); deleting is also available inside a [transaction](https://originchaindb.com/docs/schemas/transactions). SDK coverage is uneven and worth checking before you plan around it: | Operation | cURL | Python | TypeScript | Go | | --- | --- | --- | --- | --- | | Write one row | yes | db.rows.put(schema, row) | not wrapped | not wrapped | | Read one row | yes | db.rows.get(schema, pk) | not wrapped | not wrapped | | Batch write | yes | db.rows.put_batch(schema, rows) | not wrapped | not wrapped | | Delete a row | via SQL | via db.sql.execute | via db.sql | via db.SQL | | Read many rows | via SQL | db.sql.query | db.sql | db.SQL | Only the Python client wraps the row endpoints today. The TypeScript and Go examples on this page therefore call the HTTP API directly — which is exactly what those SDKs will do for you when the helpers ship. 3 ## Write and read one row. ### Write `POST /rows/:schema` takes the row as a flat JSON object — fields at the top level, not nested under a `row` key. The primary key travels in the body like any other column. A success is a terse `{ "ok": true }`: the endpoint does not echo the row back. Send an `Idempotency-Key` header on writes you might retry. A repeat of the same key replays the original response instead of writing again — which matters because a client that times out and retries has no other way to tell a lost request from a lost response. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: order-1001-v1" \ -d '{ "id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": 1714478049000 }' ``` #### Python ``` db.rows.put( "shop.orders", { "id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": 1714478049000, }, idempotency_key="order-1001-v1", # optional, but recommended ) ``` #### TypeScript ``` // The TypeScript SDK does not wrap the row endpoints yet - use fetch. const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/shop.orders`, { method: "POST", headers: { Authorization: `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", "Idempotency-Key": "order-1001-v1", }, body: JSON.stringify({ id: "ord-1001", customer: "cus-77", amount_cents: 4200, status: "pending", notes: "gift wrap", placed_ms: 1714478049000, }), }, ); if (!res.ok) throw new Error(await res.text()); ``` #### Go ``` // The Go SDK does not wrap the row endpoints yet - use net/http. body, _ := json.Marshal(map[string]any{ "id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": uint64(1714478049000), }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/rows/shop.orders", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") req.Header.Set("Idempotency-Key", "order-1001-v1") resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() ``` #### Python ``` {"ok": True} ``` #### cURL ``` { "ok": true } ``` ### Read `GET /rows/:schema/:pk` returns the row object bare — no envelope, no `rows` array. This is a hash lookup on the key, not a scan, so it does not get slower as the table grows. #### cURL ``` curl "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/ord-1001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` row = db.rows.get("shop.orders", "ord-1001") print(row["status"], row["amount_cents"]) # -> pending 4200 ``` #### TypeScript ``` const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/shop.orders/ord-1001`, { headers: { Authorization: `Bearer ${process.env.OC_TOKEN}` } }, ); if (res.status === 404) { /* not found */ } const row = await res.json(); console.log(row.status, row.amount_cents); ``` #### Go ``` req, _ := http.NewRequestWithContext(ctx, "GET", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/rows/shop.orders/ord-1001", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var row map[string]any json.NewDecoder(resp.Body).Decode(&row) fmt.Println(row["status"], row["amount_cents"]) ``` #### cURL ``` { "id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": 1714478049000, "_oc_row_version": 1 } ``` #### cURL ``` { "error": "row \"ord-1001\" not found" } ``` the _oc_row_version field Every row comes back with an engine-maintained `_oc_row_version` counter — 1 on first write, incremented on every subsequent one. It is useful for spotting that a row changed. It is not a concurrency control: no write path checks it, and there is no `If-Match` or expected-version parameter on any HTTP write. Do not write it back — strip it before re-posting a row you read. 4 ## Upsert semantics — read this one. the row endpoint is an upsert, not an insert `POST /rows/:schema` with a primary key that already exists replaces the existing row and returns 200. It is last-write-wins by design. There is no duplicate-key error on this path. If you need "fail if it already exists", use SQL `INSERT` — [below](https://originchaindb.com/docs/schemas/row-crud#insert-strict). This is deliberate. The underlying store is key-addressed, so writing a key overwrites it, and the row endpoint keeps that contract untouched because it is also the high-throughput ingest path. What the engine does do on an overwrite is retire the old row's secondary-index and relation entries in the same atomic write, so an overwrite leaves nothing stale behind. SQL `INSERT` is strict. A duplicate primary key over `/sql` returns `409 constraint_violation` and writes nothing — the existing row is preserved. It catches duplicates against committed state and duplicates within the same multi-row `VALUES` list. Use `INSERT … ON CONFLICT` when you want an explicit, controlled upsert. ?expect=insert does not do what its name suggests The `?expect=insert` query parameter (and the Python client's `expect_insert=True`) is a performance hint, not an assertion. It tells the engine to skip reading the prior row, which saves a lookup per row on bulk ingest. It does not make a duplicate key fail — the write still overwrites — and because the prior row was never read, its old secondary-index entries are not retired, leaving stale index entries behind. Use it only when you know the keys are new. 5 ## Updating a row. every row write is a full replace There is no merge or partial-update route. Whatever object you post becomes the row — any column you leave out is erased, not preserved. This is the single most common way to lose data through this API. Two ways to change one field. Prefer the SQL form: it does the read-modify-write inside the engine, so there is no window between your read and your write, and no chance of dropping a column you forgot to carry over. #### cURL ``` # There is no PATCH. Re-POST the WHOLE row - every field you omit # is erased. Read first, mutate, write back. ROW=$(curl -s "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/ord-1001" \ -H "Authorization: Bearer $OC_TOKEN") echo "$ROW" \ | jq 'del(._oc_row_version) | .status = "shipped"' \ | curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" --data-binary @- # Or let the engine do the read-modify-write for you, with SQL: curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{"sql":"UPDATE shop.orders SET status = $1 WHERE id = $2", "params":["shipped","ord-1001"]}' ``` #### Python ``` # Option A - full-row replace through the row endpoint. row = db.rows.get("shop.orders", "ord-1001") row.pop("_oc_row_version", None) # engine-maintained; do not write it back row["status"] = "shipped" db.rows.put("shop.orders", row) # Option B - let the engine do the read-modify-write. res = db.sql.execute("UPDATE shop.orders SET status = 'shipped' WHERE id = 'ord-1001'") print(res.kind) # -> update ``` #### TypeScript ``` // Option A - full-row replace (no SDK row methods yet, so fetch). const url = `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/shop.orders`; const H = { Authorization: `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json" }; const row = await (await fetch(`${url}/ord-1001`, { headers: H })).json(); delete row._oc_row_version; row.status = "shipped"; await fetch(url, { method: "POST", headers: H, body: JSON.stringify(row) }); // Option B - the SDK does wrap SQL. const res = await db.sql( "UPDATE shop.orders SET status = 'shipped' WHERE id = 'ord-1001'", ); ``` #### Go ``` // Option A - full-row replace (no SDK row methods yet, so net/http). base := "https://" + os.Getenv("OC_HOST") + "/v1/tenants/" + os.Getenv("OC_TENANT") req, _ := http.NewRequestWithContext(ctx, "GET", base+"/rows/shop.orders/ord-1001", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, _ := http.DefaultClient.Do(req) var row map[string]any json.NewDecoder(resp.Body).Decode(&row) resp.Body.Close() delete(row, "_oc_row_version") row["status"] = "shipped" body, _ := json.Marshal(row) req, _ = http.NewRequestWithContext(ctx, "POST", base+"/rows/shop.orders", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, _ = http.DefaultClient.Do(req) resp.Body.Close() // Option B - the SDK does wrap SQL. _, err := db.SQL(ctx, "UPDATE shop.orders SET status = 'shipped' WHERE id = 'ord-1001'") ``` #### cURL ``` { "kind": "update", "schema": "shop.orders", "rows_affected": 1 } ``` Two rules on the SQL form: a `WHERE` clause is mandatory — a bare `UPDATE` is refused as a safety check — and you cannot `SET` a primary-key column. To change a key, write the new row and delete the old one. 6 ## Deleting a row. There is no `DELETE /rows/:schema/:pk` route. Deleting goes through SQL, or through the transaction row surface where the verb does exist. The delete is real, not a tombstone: the row body, every secondary-index entry and every relation entry derived from it are removed in one atomic write, and a subsequent point read returns `404`. #### cURL ``` # There is NO DELETE /rows/:schema/:pk route. Use SQL: curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{"sql":"DELETE FROM shop.orders WHERE id = $1","params":["ord-1001"]}' # { "kind": "delete", "schema": "shop.orders", # "pk": "ord-1001", "rows_affected": 1 } # ...or delete inside a transaction, where the verb IS available: curl -X DELETE \ "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/rows/shop.orders/ord-1001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` # The Python SDK has no rows.delete - go through SQL. res = db.sql.execute("DELETE FROM shop.orders WHERE id = 'ord-1001'") print(res.kind) # -> delete ``` #### TypeScript ``` const res = await db.sql("DELETE FROM shop.orders WHERE id = 'ord-1001'"); if (res.kind === "delete") console.log(res.schema, res.pk); ``` #### Go ``` res, err := db.SQL(ctx, "DELETE FROM shop.orders WHERE id = 'ord-1001'") if err != nil { /* handle */ } fmt.Println(res.Kind, res.Schema, res.PK) // -> delete shop.orders ord-1001 ``` deleting a row that is not there It is a silent success, not a `404`. Nothing is written and the call returns normally. If you need to know whether a row existed, read it first, or use `DELETE … RETURNING` over SQL and check whether any rows came back. 7 ## Batch and streaming writes. `POST /rows/:schema/_batch` is the bulk path, and the throughput difference is not marginal — the batch route sustains roughly an order of magnitude more rows per second than issuing the same rows one request at a time, because the whole array is prepared once and committed in a single durable write. The body is a raw array of row objects. Wrapping it as `{ "rows": [...] }` is the most common mistake here and returns a `400` about expecting a sequence. #### cURL ``` # The body is a RAW ARRAY. Not { "rows": [...] }. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/_batch" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '[ { "id": "ord-2001", "customer": "cus-12", "amount_cents": 1500, "status": "paid", "notes": "", "placed_ms": 1714478100000 }, { "id": "ord-2002", "customer": "cus-12", "amount_cents": 8900, "status": "paid", "notes": "fragile", "placed_ms": 1714478160000 }, { "id": "ord-2003", "customer": "cus-31", "amount_cents": 400, "status": "pending", "notes": "", "placed_ms": 1714478220000 } ]' ``` #### Python ``` rows = [ {"id": "ord-2001", "customer": "cus-12", "amount_cents": 1500, "status": "paid", "notes": "", "placed_ms": 1714478100000}, {"id": "ord-2002", "customer": "cus-12", "amount_cents": 8900, "status": "paid", "notes": "fragile", "placed_ms": 1714478160000}, {"id": "ord-2003", "customer": "cus-31", "amount_cents": 400, "status": "pending", "notes": "", "placed_ms": 1714478220000}, ] # put_batch slices into chunks of 1000 and posts each one. n = db.rows.put_batch("shop.orders", rows) print(n) # -> 3 (total inserted across all chunks) ``` #### TypeScript ``` // No SDK batch method yet - post the raw array with fetch. const rows = [ { id: "ord-2001", customer: "cus-12", amount_cents: 1500, status: "paid", notes: "", placed_ms: 1714478100000 }, { id: "ord-2002", customer: "cus-12", amount_cents: 8900, status: "paid", notes: "fragile", placed_ms: 1714478160000 }, { id: "ord-2003", customer: "cus-31", amount_cents: 400, status: "pending", notes: "", placed_ms: 1714478220000 }, ]; const res = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/rows/shop.orders/_batch`, { method: "POST", headers: { Authorization: `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify(rows), // a raw array, NOT { rows } }, ); const { inserted } = await res.json(); ``` #### Go ``` // No SDK batch method yet - post the raw array with net/http. rows := []map[string]any{ {"id": "ord-2001", "customer": "cus-12", "amount_cents": 1500, "status": "paid", "notes": "", "placed_ms": uint64(1714478100000)}, {"id": "ord-2002", "customer": "cus-12", "amount_cents": 8900, "status": "paid", "notes": "fragile", "placed_ms": uint64(1714478160000)}, {"id": "ord-2003", "customer": "cus-31", "amount_cents": 400, "status": "pending", "notes": "", "placed_ms": uint64(1714478220000)}, } body, _ := json.Marshal(rows) // a raw array, NOT {"rows": [...]} req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/rows/shop.orders/_batch", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var out struct{ Inserted int `json:"inserted"` } json.NewDecoder(resp.Body).Decode(&out) ``` #### cURL ``` { "inserted": 3 } ``` the JSON array form is all-or-nothing One array is one log frame. If any row fails validation or a constraint, the request fails and none of the rows are written. That is a real guarantee you can lean on for a bounded batch — and it is exactly why this form is capped at 64 MiB rather than being unbounded. ### Streaming ingest for anything larger Send `Content-Type: application/x-ndjson` and the same route switches to a streaming loader: one JSON object per line, flushed in chunks, with no cap on the total body. This is the path for a multi-gigabyte load. #### cURL ``` # One JSON object per line. Nothing is held whole in memory, so this # is the path for files that do not fit in a request body. curl -X POST \ "https://$OC_HOST/v1/tenants/$OC_TENANT/rows/shop.orders/_batch?chunk=2000" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/x-ndjson" \ --data-binary @orders.ndjson ``` #### cURL ``` { "inserted": 480000 } ``` #### cURL ``` { "inserted": 312000, "error": "..." } ``` streaming ingest is NOT all-or-nothing Chunks commit as they go. A failure partway returns `207 Multi-Status` with an `inserted` count and an `error` — and that count is real: those rows are already durable. Treat `207` as a partial success and resume from the reported offset. Note also that a `207` is not cached against your idempotency key, so a blind retry re-sends everything. Tune the flush size with `?chunk=N`. The default is 1,000 rows per flush and values above 10,000 are clamped down silently rather than rejected. Larger chunks trade memory for fewer flushes. 8 ## Reading many rows. There is no "list rows" or "scan table" route on the row surface, and no cursor. To read a set of rows, use SQL — which every SDK wraps, and which lets you project only the columns you need. #### cURL ``` # There is no "list rows" route. Read sets of rows with SQL. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{"sql":"SELECT id, customer, amount_cents FROM shop.orders WHERE status = $1 ORDER BY placed_ms DESC LIMIT 50", "params":["paid"]}' ``` #### Python ``` result = db.sql.query( """SELECT id, customer, amount_cents FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 50""" ) for row in result.rows: print(row["id"], row["amount_cents"]) ``` #### TypeScript ``` const res = await db.sql(` SELECT id, customer, amount_cents FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 50 `); if (res.kind === "select") { for (const row of res.rows) console.log(row); } ``` #### Go ``` res, err := db.SQL(ctx, ` SELECT id, customer, amount_cents FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 50 `) if err != nil { /* handle */ } for _, row := range res.Rows { fmt.Println(row["id"], row["amount_cents"]) } ``` always put a LIMIT on it A projection over a large table is assembled in memory before it is sent, and past the per-query cap the request fails with `413 result_too_large` rather than streaming. Keep a `LIMIT` on anything that reads a big table, and page with `OFFSET` — or aggregate server-side, which does stream. See [the SQL limits](https://originchaindb.com/docs/sql#limits) for the details. 9 ## What gets validated. Every write path — single row, batch array, streaming ingest, SQL, and transactions — goes through the same validation. What it checks: - Required columns are present. A missing column declared `required = true` is a `400`. An explicit `null` on a required column is the same error. - Types match, they are not coerced. A string where an `i64` is declared is rejected, not parsed. - Some values are canonicalised. A decimal sent as a JSON number is stored as its exact decimal string; a date or timestamp sent as a text literal is stored as its integer form. Both spellings address the same row, including in the primary key. - CHECK constraints and unique indexes are enforced on every write, including the bulk path. A violation is a `409`. - Foreign keys are enforced before the write is queued. An orphan reference is a `409`. undeclared columns are accepted, not rejected Validation walks the columns your manifest declares — it never walks the keys of your request object. A field that is not in the schema is stored verbatim and returned on read. There is no strict-schema mode. That makes ad-hoc extra fields easy, and it makes a typo like `amount_cent` silently create a junk field while the real column stays unset — which then trips its own required-column or CHECK error, if you declared one. Declare `required = true` on the columns you cannot do without. | HTTP | When | Body | | --- | --- | --- | | 400 | A required column is missing, a value has the wrong JSON type, or the primary key in the path will not coerce to the column's type. | { "error": "…" } | | 400 | A point read against a table whose primary key spans more than one column. | composite primary keys are not addressable by path… | | 404 | The row does not exist, or the table is not registered. | { "error": "row \"…\" not found" } | | 409 | A foreign key, CHECK or UNIQUE-index constraint failed. | { "error": "constraint_violation", "detail": "…" } | | 413 | A single-row body over 8 MiB, a batch body over 64 MiB, or a newline-delimited line over 1 MiB. | { "error": "body too large or unreadable: …" } | | 429 | The write queue is saturated. Carries Retry-After: 1. | { "error": "write queue full (n commits pending); retry" } | | 501 | On a sharded instance, a foreign key pointing at a table owned by another shard. | { "error": "…" } | 10 ## Limits & gotchas. | Limit | Value | Notes | | --- | --- | --- | | Single-row body | 8 MiB | The general request limit. A row larger than this must go through the batch route. | | Batch body (JSON array) | 64 MiB | Eight times the single-row limit, because the whole array is buffered to commit atomically. | | Rows per batch | no cap | There is no row-count limit — the byte size is the real constraint. | | Streaming body | unlimited | Per line, 1 MiB. In-flight buffer, 64 MiB — lower `?chunk=N` if you hit it. | | Streaming flush size | 1,000 / 10,000 | Default and maximum `?chunk=N`. A larger value is clamped, not refused. | | Write queue depth | 4,096 | Pending commits. Over it, writes shed with `429` and `Retry-After: 1` — back off rather than hammering. | no optimistic concurrency on this surface No `If-Match`, no expected-version parameter, no compare-and-set. Two clients doing read-modify-write on the same row through the row endpoint will both succeed and one update is lost, silently. If that matters, use a [transaction](https://originchaindb.com/docs/schemas/transactions), whose commit revalidates what you read. composite primary keys have no point-read route `GET /rows/:schema/:pk` addresses one path segment. Against a table whose primary key spans two or more columns it returns `400`, naming the number of key columns. Read those tables with SQL: `WHERE k1 = … AND k2 = …`. Writing is unaffected — the key columns travel in the body like any other. a degraded write is still a successful write Replication to a standby is asynchronous: the `200` means the write is flushed to durable storage on this instance, not that the standby holds it. Where a standby-acknowledgement wait is configured and the standby does not answer in time, the write still returns `200` and the response carries `X-OC-Replication: degraded`. If your application cares about that distinction, check the header; a bare status code will not tell you. Row writes are the only endpoints that wait at all — SQL, Cypher and transaction commits never do. a single write is not slower than a batched one, per request Single-row writes go through the same group-commit path as batches — concurrent writes are coalesced into one flush rather than each paying for their own. So a busy single-row workload scales fine. The batch route wins on bulk loading, where you have the rows in hand and can skip the per-request overhead entirely. ## Related. - [SQL](https://originchaindb.com/docs/schemas/sql) — reading sets of rows, partial updates, deletes, and strict inserts. - [Transactions](https://originchaindb.com/docs/schemas/transactions) — making several row writes land together, with conflict detection. - [Schema reference](https://originchaindb.com/docs/schemas/reference) — every field in the manifest, including column types and constraints. - [Error reference](https://originchaindb.com/docs/errors) — the shared error envelope every endpoint returns. --- # Schema for SQL - OriginChainDB Canonical source: https://originchaindb.com/docs/schemas/sql Sitemap last modified: 2026-09-13T20:18:50.000Z schema · sql # SQL. 1 ## Before you start. This page uses the `shop.orders` table from the [quickstart](https://originchaindb.com/docs/quickstart). Table names are always fully qualified: `namespace.table`, never a bare table name. ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # Turns WHERE status = '...' into a sub-linear lookup. [[indexes]] name = "by_status" columns = ["status"] ``` The `[[indexes]]` block is what turns `WHERE status = …` from a full scan into a lookup. [EXPLAIN](https://originchaindb.com/docs/sql#ex-explain) tells you which one you got. 2 ## Type mapping. What a column declared with each `ty` looks like coming out of a `SELECT`, and what you get after decoding in each client. All three clients decode rows into a generic map, so the language column is what JSON decoding yields there. | ty | JSON on the wire | TypeScript | Python | Go | | --- | --- | --- | --- | --- | | i64 / u64 | number | number | int | float64 ⚠ | | f64 | number | number | float | float64 | | bool | boolean | boolean | bool | bool | | str / text | string | string | str | string | | decimal | string — "19.99" | string | str | string | | uuid | string | string | str | string | | enum | string | string | str | string | | inet | string | string | str | string | | bytes | hex string — "\\x4f43" | string | str | string | | date | text — "2026-07-25" | string | str | string | | timestamp | text — "2026-07-25 09:14:22" | string | str | string | | time | text — "09:14:22.500000" | string | str | string | | interval | number — milliseconds | number | int | float64 | | json | object or array | unknown | dict / list | map[string]any / []any | | list | array | unknown[] | list | []any | | point | { lat, lng } object | object | dict | map[string]any | decimal is a string, and its arithmetic is exact A `decimal` column travels as a JSON string — `"19.99"` — deliberately, so no binary-float rounding is ever introduced into money. On write you may send either a string or a bare JSON number; both are canonicalised to the same stored form. `SUM`, `AVG`, arithmetic, `ORDER BY` and range comparisons over a decimal column are all exact and numeric, not lexical, and the result comes back in the same canonical string form. A result past the exact-decimal range is an error, never a silently wrong number. large integers lose precision in TypeScript and Go An `i64` or `u64` travels as a JSON number. JavaScript decodes it to a double, and Go decodes into `map[string]any` as `float64` — so values beyond about 9×10¹⁵ are silently rounded in both. Python is unaffected. If you store identifiers or counters that big, declare the column as `str`, or decode the response yourself with a big-integer-aware parser. dates and times come back as text `date`, `time` and `timestamp` columns are stored as integers but rendered into familiar text on the way out — you read `"2026-07-25"`, not an epoch count. That reformatting applies only to output columns the plan can prove are a straight pass-through of a temporal column, so a computed expression over one still comes back as a number. On write, both spellings are accepted and address the same row. ## Related. - [Row CRUD](https://originchaindb.com/docs/schemas/row-crud) — point reads and the bulk-write path, and how its upsert differs from SQL `INSERT`. - [Transactions](https://originchaindb.com/docs/schemas/transactions) — making several statements land together, and what isolation you actually get. - [Schema reference](https://originchaindb.com/docs/schemas/reference) — column types, indexes, foreign keys and CHECK constraints. - [Queries in the dashboard](https://originchaindb.com/docs/dashboard/queries) — running the same statements from the console workbench. - [Error reference](https://originchaindb.com/docs/errors) — the shared error envelope every endpoint returns. --- # Transactions and snapshot isolation - OriginChainDB Canonical source: https://originchaindb.com/docs/schemas/transactions Sitemap last modified: 2026-09-13T20:18:50.000Z query shapes · transactions # Transactions. A transaction is how you make several writes land together or not at all. Reach for one when a single logical change touches more than one row — moving stock between two SKUs, debiting one balance to credit another, writing a record plus its audit trail. If your change is one row, you do not need a transaction: a single row write is already atomic. OriginChainDB ships two transaction surfaces with different guarantees. Picking the wrong one is a correctness bug, so the difference is the first thing on this page. preview surface Every `/tx` response carries `"_preview": true`. The wire shape can still change between releases — pin your client and read the changelog before upgrading. No OriginChainDB SDK wraps transactions yet, in any language, so every example below drives the endpoints over plain HTTP. 1 ## Which surface to use. | | /tx endpoints | SQL session transaction | | --- | --- | --- | | Isolation | Snapshot isolation by default; `serializable` opt-in. | Read committed. Not selectable. | | Conflict detection | Yes — read set revalidated at commit, 409 on loss. | None. Concurrent writers do not abort each other. | | Work you can do | Row put / get / delete, plus scan-based read plans. | Any SQL write statement, including data-definition statements. | | Handle | `tx_id` in the URL path. | `X-OC-Session-Id` header. | | Reach for it when | The change is read-modify-write and a lost update would be a bug. | You want several SQL statements to land in one frame and there are no concurrent writers to the same rows. | Sections 3 to 6 cover the `/tx` endpoints. [Section 7](https://originchaindb.com/docs/schemas/transactions#sqltx) covers SQL session transactions. 2 ## What the isolation actually guarantees. The default level on `/tx` is snapshot isolation, implemented as optimistic, commit-time validation. The mechanics are worth understanding, because one detail surprises people: - Reads hit live data, not a frozen snapshot. There are no version chains. Each read returns whatever the store holds at that instant and records the value it saw in a read set. - The snapshot is enforced at commit, not at read. Commit re-reads every key in the read set and compares. If anything changed, the commit aborts with `409 tx_conflict` and nothing is written. - So a committed transaction saw a consistent view — but a re-read mid-transaction is not repeatable. Read the same key twice and the second read can return a newer value. The commit will then fail, so you cannot act on the inconsistency, but do not write code that assumes two reads agree. - Writes are buffered. Nothing is durable until commit, at which point the whole buffer lands in a single write-ahead-log frame — the storage layer's all-or-nothing unit. - You see your own writes. A read through the transaction returns its own buffered value; a buffered delete reads as absent. snapshot isolation is not serializable At the default level, phantoms and write skew are possible. A scan that matched three rows at the start can match four at commit if another writer inserted one, and the conflict check will not catch it — it validates the keys you observed, not the keys that would now match. Two transactions that read a shared invariant and each write a different row can both commit and break it. If your correctness argument depends on either of those not happening, use [serializable](https://originchaindb.com/docs/schemas/transactions#serializable). Which reads count toward that check depends on how you read. This table is the most important thing on the page: | How you read | Snapshot isolation | Serializable | | --- | --- | --- | | GET /tx/:id/rows/:schema/:pk | Recorded. A concurrent change to that row aborts your commit with tx_conflict. | Recorded, plus a read predicate for the cycle check. | | POST /tx/:id/query | NOT recorded. The scan gives you read-your-writes but no conflict protection at all. | Recorded as read predicates — this is what catches phantoms and write skew. | | POST /sql or GET /rows (outside the transaction) | Not recorded. Invisible to the transaction. | Not recorded. Invisible to the transaction. | the sharpest edge on this page At the default level, a transaction that decides what to write based on a `/tx/:id/query` scan gets no conflict protection for that decision. The scan reflects your own buffered writes, but it records nothing in the read set, so the commit has nothing to revalidate and will succeed even if the scanned rows changed underneath you. Either read the specific keys you care about with `GET /tx/:id/rows/…` so they enter the read set, or open the transaction as [serializable](https://originchaindb.com/docs/schemas/transactions#serializable). also not present No range or gap locking — scans do not lock what they scanned. No deadlock or starvation defence: a long transaction contesting a hot key can keep losing the race, so bound your retries. No savepoints and no nested transactions. No data-definition statements inside a `/tx` transaction — schema migrations run through their own cutover path. Conflicts abort at commit and never block, so there are no lock-wait deadlocks by construction. 3 ## Before you start. Transactions add no fields to your schema — they wrap writes against tables you have already registered. Every example below uses the `shop.orders` table from the [quickstart](https://originchaindb.com/docs/quickstart): ``` namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # Enforced INSIDE a transaction too - a violation aborts the whole # transaction at COMMIT, not just the offending write. [[check_constraints]] name = "amount_positive" expression = "amount_cents > 0" ``` Constraints declared on the table — foreign keys, back-references, `CHECK`, uniqueness — are enforced at commit, on the whole buffer. One bad row aborts the entire transaction, not just that write. The shell examples assume `$OC_HOST`, `$OC_TENANT` and `$OC_TOKEN` are set — see [Authentication](https://originchaindb.com/docs/auth). 4 ## The lifecycle, end to end. Five endpoints, all under `/v1/tenants/:tenant/tx`. Open one, buffer work against its id, then commit or roll back. ### Open a transaction `POST /tx/begin` takes no body. It returns a ULID `tx_id`, the logical read tick the transaction is anchored to, and the isolation level you actually got — always check that field rather than assuming. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/begin" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` import os, requests BASE = f"https://{os.environ['OC_HOST']}/v1/tenants/{os.environ['OC_TENANT']}" H = {"Authorization": f"Bearer {os.environ['OC_TOKEN']}"} # No SDK method wraps /tx yet - plain HTTP. tx = requests.post(f"{BASE}/tx/begin", headers=H).json() tx_id = tx["tx_id"] print(tx["isolation"]) # -> snapshot_isolation ``` #### TypeScript ``` const BASE = `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}`; const H = { Authorization: `Bearer ${process.env.OC_TOKEN}` }; // No SDK method wraps /tx yet - plain fetch. const res = await fetch(`${BASE}/tx/begin`, { method: "POST", headers: H }); const tx = await res.json(); const txId: string = tx.tx_id; console.log(tx.isolation); // -> snapshot_isolation ``` #### Go ``` base := "https://" + os.Getenv("OC_HOST") + "/v1/tenants/" + os.Getenv("OC_TENANT") // No SDK method wraps /tx yet - plain net/http. req, _ := http.NewRequestWithContext(ctx, "POST", base+"/tx/begin", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } defer resp.Body.Close() var tx struct { TxID string `json:"tx_id"` ReadTS uint64 `json:"read_ts"` Isolation string `json:"isolation"` } json.NewDecoder(resp.Body).Decode(&tx) fmt.Println(tx.Isolation) // -> snapshot_isolation ``` #### cURL ``` { "tx_id": "01JTS0Q4M9V3K7X2N8R1P6C0DE", "read_ts": 184213, "isolation": "snapshot_isolation", "_preview": true } ``` ### Buffer writes `POST /tx/:tx_id/rows/:schema` buffers an upsert. The row body is flat — the same shape the non-transactional row endpoint takes. It is validated against the manifest immediately, so a type or shape error comes back as a `400` right here rather than at commit. Referential and `CHECK` constraints are the ones deferred to commit. #### cURL ``` # Two writes, one transaction. Neither is durable yet. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/rows/shop.orders" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{ "id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": 1714478049000 }' # { "tx_id": "01JTS0Q4M9V3K7X2N8R1P6C0DE", "op": "put", # "buffered": true, "_preview": true } curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/rows/shop.orders" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{ "id": "ord-1002", "customer": "cus-77", "amount_cents": 1500, "status": "pending", "notes": "", "placed_ms": 1714478051000 }' ``` #### Python ``` rows = [ {"id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": 1714478049000}, {"id": "ord-1002", "customer": "cus-77", "amount_cents": 1500, "status": "pending", "notes": "", "placed_ms": 1714478051000}, ] for row in rows: r = requests.post(f"{BASE}/tx/{tx_id}/rows/shop.orders", headers=H, json=row) r.raise_for_status() # 400 here = the row failed manifest validation assert r.json()["buffered"] is True ``` #### TypeScript ``` const rows = [ { id: "ord-1001", customer: "cus-77", amount_cents: 4200, status: "pending", notes: "gift wrap", placed_ms: 1714478049000 }, { id: "ord-1002", customer: "cus-77", amount_cents: 1500, status: "pending", notes: "", placed_ms: 1714478051000 }, ]; for (const row of rows) { const r = await fetch(`${BASE}/tx/${txId}/rows/shop.orders`, { method: "POST", headers: { ...H, "Content-Type": "application/json" }, body: JSON.stringify(row), }); if (!r.ok) throw new Error(await r.text()); // 400 = manifest validation } ``` #### Go ``` rows := []map[string]any{ {"id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": uint64(1714478049000)}, {"id": "ord-1002", "customer": "cus-77", "amount_cents": 1500, "status": "pending", "notes": "", "placed_ms": uint64(1714478051000)}, } for _, row := range rows { body, _ := json.Marshal(row) req, _ := http.NewRequestWithContext(ctx, "POST", base+"/tx/"+txID+"/rows/shop.orders", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) if err != nil { /* handle */ } resp.Body.Close() } ``` ### Read through the transaction `GET /tx/:tx_id/rows/:schema/:pk` answers from the transaction's own buffer first, then from committed state. The read is recorded in the read set — which is exactly what makes a concurrent change to that row abort this transaction at commit. #### cURL ``` # Reads the tx's own buffered write first, then committed state. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/rows/shop.orders/ord-1001" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` body = requests.get( f"{BASE}/tx/{tx_id}/rows/shop.orders/ord-1001", headers=H ).json() row = body["row"] # note the "row" envelope print(row["amount_cents"]) # -> 4200, from this tx's own buffer ``` #### TypeScript ``` const body = await ( await fetch(`${BASE}/tx/${txId}/rows/shop.orders/ord-1001`, { headers: H }) ).json(); const row = body.row; // note the "row" envelope console.log(row.amount_cents); // -> 4200, from this tx's own buffer ``` #### Go ``` req, _ = http.NewRequestWithContext(ctx, "GET", base+"/tx/"+txID+"/rows/shop.orders/ord-1001", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, _ = http.DefaultClient.Do(req) defer resp.Body.Close() var body struct { Row map[string]any `json:"row"` // note the "row" envelope } json.NewDecoder(resp.Body).Decode(&body) fmt.Println(body.Row["amount_cents"]) // -> 4200, from this tx's own buffer ``` #### cURL ``` { "row": { "id": "ord-1001", "customer": "cus-77", "amount_cents": 4200, "status": "pending", "notes": "gift wrap", "placed_ms": 1714478049000, "_oc_row_version": 1 }, "tx_id": "01JTS0Q4M9V3K7X2N8R1P6C0DE", "_preview": true } ``` ### Buffer a delete `DELETE /tx/:tx_id/rows/:schema/:pk` buffers a removal. Note the commit-time cost: any transaction with a buffered delete makes the engine inspect every table on the instance for references back to the deleted rows, where a write-only transaction inspects just the tables it touched. #### cURL ``` curl -X DELETE \ "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/rows/shop.orders/ord-0900" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` requests.delete( f"{BASE}/tx/{tx_id}/rows/shop.orders/ord-0900", headers=H ).raise_for_status() ``` #### TypeScript ``` await fetch(`${BASE}/tx/${txId}/rows/shop.orders/ord-0900`, { method: "DELETE", headers: H, }); ``` #### Go ``` req, _ = http.NewRequestWithContext(ctx, "DELETE", base+"/tx/"+txID+"/rows/shop.orders/ord-0900", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, _ = http.DefaultClient.Do(req) resp.Body.Close() ``` #### cURL ``` { "tx_id": "01JTS0Q4M9V3K7X2N8R1P6C0DE", "op": "delete", "buffered": true, "_preview": true } ``` ### Commit `POST /tx/:tx_id/commit` revalidates the read set, runs the constraint checks, and writes the buffer as one frame. On success the body carries the log position the frame landed at. On failure nothing was written, and the `error` field tells you whether retrying can help. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/commit" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` r = requests.post(f"{BASE}/tx/{tx_id}/commit", headers=H) if r.status_code == 200: print("committed at", r.json()["lsn"]) elif r.status_code == 409: err = r.json() # "tx_conflict" | "serialization_failure" -> retry the whole tx. # "constraint_violation" -> fix the data first. print("aborted:", err.get("error")) ``` #### TypeScript ``` const r = await fetch(`${BASE}/tx/${txId}/commit`, { method: "POST", headers: H }); if (r.status === 200) { console.log("committed at", (await r.json()).lsn); } else if (r.status === 409) { const err = await r.json(); // "tx_conflict" | "serialization_failure" -> retry the whole tx. // "constraint_violation" -> fix the data first. console.log("aborted:", err.error); } ``` #### Go ``` req, _ = http.NewRequestWithContext(ctx, "POST", base+"/tx/"+txID+"/commit", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, _ = http.DefaultClient.Do(req) defer resp.Body.Close() switch resp.StatusCode { case 200: var ok struct{ Committed bool } json.NewDecoder(resp.Body).Decode(&ok) case 409: var e struct{ Error string } json.NewDecoder(resp.Body).Decode(&e) // "tx_conflict" | "serialization_failure" -> retry the whole tx. // "constraint_violation" -> fix the data first. } ``` #### cURL ``` { "committed": true, "tx_id": "01JTS0Q4M9V3K7X2N8R1P6C0DE", "lsn": { "segment": 12, "offset": 884210 }, "_preview": true } ``` #### cURL ``` { "error": "tx_conflict", "tx_id": "01JTS0...", "read_lsn": 184213, "current_lsn": 184298, "_preview": true } ``` about the lsn field On a single-writer instance `lsn` is an object — `{ segment, offset }`. On an instance running consensus replication the position is stamped after the round, and the field comes back `null`; the response header `X-OC-Replication-Path` tells you which path served you. A commit that buffered nothing also succeeds, and reports the sentinel `{ segment: 0, offset: 0 }` — nothing was appended. Treat `committed: true` as the durability signal, not the shape of `lsn`. commit always forces a disk flush A `/tx` commit flushes its log frame to stable storage before it answers, regardless of how the instance's general write-flush policy is tuned. A crash immediately after a `200` cannot lose the transaction. That also means a commit is more expensive than an ordinary row write — batch your work into fewer, larger transactions rather than many tiny ones. ### Roll back, and check state `POST /tx/:tx_id/rollback` discards the buffer. `GET /tx/:tx_id` reports the current state and survives the terminal transition, so you can always ask what happened to an id. #### cURL ``` # Discard everything the transaction buffered. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/rollback" \ -H "Authorization: Bearer $OC_TOKEN" # Ask what happened to a transaction at any time. curl "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### Python ``` requests.post(f"{BASE}/tx/{tx_id}/rollback", headers=H) state = requests.get(f"{BASE}/tx/{tx_id}", headers=H).json()["state"] # active | committing | committed | rolled_back | conflict | expired ``` #### TypeScript ``` await fetch(`${BASE}/tx/${txId}/rollback`, { method: "POST", headers: H }); const { state } = await (await fetch(`${BASE}/tx/${txId}`, { headers: H })).json(); // active | committing | committed | rolled_back | conflict | expired ``` #### Go ``` req, _ = http.NewRequestWithContext(ctx, "POST", base+"/tx/"+txID+"/rollback", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, _ = http.DefaultClient.Do(req) resp.Body.Close() req, _ = http.NewRequestWithContext(ctx, "GET", base+"/tx/"+txID, nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, _ = http.DefaultClient.Do(req) defer resp.Body.Close() var st struct{ State string } json.NewDecoder(resp.Body).Decode(&st) // active | committing | committed | rolled_back | conflict | expired ``` #### cURL ``` { "rolled_back": true, "tx_id": "01JTS0...", "_preview": true } ``` #### cURL ``` { "tx_id": "01JTS0...", "read_ts": 184213, "state": "committed", "isolation": "snapshot_isolation", "_preview": true } ``` | state | Meaning | | --- | --- | | `active` | Open. Accepts row ops, queries, commit and rollback. | | `committing` | A commit is in flight. Transient. | | `committed` | Terminal. Every buffered op landed in one write-ahead-log frame. | | `rolled_back` | Terminal. Buffer discarded — by an explicit rollback, or by a commit that failed a constraint check or an internal error. | | `conflict` | Terminal. Commit lost the race: a row in the read set changed, or (serializable only) the commit-time cycle check fired. | | `expired` | Terminal. The idle sweeper aborted it after no client activity for the idle window. Buffer discarded, exactly as on rollback. | Rolling back a transaction that is already terminal returns `409`, not a silent success. An unknown or other-instance `tx_id` returns `404` — deliberately indistinguishable, so an id cannot be probed across instances. Terminal handles are kept for about an hour so a late status poll still gets a truthful answer; after that the same poll returns `404`. 5 ## Serializable isolation. Pass `?isolation=serializable` to `/tx/begin` and the engine layers serializable snapshot isolation over the machinery above: reads register predicates, writes register predicates for both the before and after image, and commit runs a dangerous-structure check that aborts the pivot of a read-write dependency cycle. Write skew and phantoms are rejected, not merely observed, and the failure arrives as `409 serialization_failure`. The only other accepted value is `snapshot_isolation` (the default). Anything else is a `400`. #### cURL ``` # Opt in at begin. The response echoes what you actually got. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/begin?isolation=serializable" \ -H "Authorization: Bearer $OC_TOKEN" ``` #### cURL ``` { "tx_id": "01JTS1...", "read_ts": 184301, "isolation": "serializable", "_preview": true } ``` #### cURL ``` { "error": "serializable_sharded_unsupported", "isolation": "serializable", "available_isolation": ["snapshot_isolation"] } ``` single-shard only The dependency graph that makes serializable work is held in memory inside a single serving process. On a sharded instance it cannot see a peer's dependencies, so `begin` refuses the request with `501 serializable_sharded_unsupported` and lists `available_isolation`. This is deliberate: you learn the level is unavailable before you build a transaction on a guarantee you would not have got. A transaction opened at single-shard whose instance is resharded mid-flight is refused at its next query or at commit for the same reason. reads must go through the transaction Serializable only sees the reads you make through the transaction — `GET /tx/:id/rows/…` and `POST /tx/:id/query`. A read issued on the ordinary `/sql` or `/rows` endpoints registers nothing, so basing a serializable write on it gives you no guarantee at all. single active writer, too The same in-memory constraint applies to consensus replication. On an instance running quorum replication serializable is accepted for point-read transactions and enforced by re-validating the read set at commit; a scan or range read inside a serializable transaction is refused with `501 serializable_scan_raft_unsupported`. Serializable stays single-shard only. Snapshot isolation works everywhere. 6 ## Retrying a conflict. Optimistic concurrency means conflicts are normal, not exceptional. Any client that writes contended rows needs a retry loop. Two rules: - Retry the whole transaction, from a fresh begin. A conflicted transaction is terminal — you cannot resume it, and reusing the id returns `409`. Re-read your inputs too; the values you decided on are the ones that just went stale. - Bound the loop and back off. There is no starvation defence in the engine, so an unbounded retry against a hot key is an infinite loop you wrote yourself. Retry on `tx_conflict` and `serialization_failure`. Do not retry `constraint_violation` — the same data will fail again. #### Python ``` import time, requests RETRYABLE = {"tx_conflict", "serialization_failure"} def run_tx(apply, attempts=5): """apply(tx_id) buffers the work. Returns the commit body.""" for n in range(attempts): tx_id = requests.post(f"{BASE}/tx/begin", headers=H).json()["tx_id"] try: apply(tx_id) except Exception: requests.post(f"{BASE}/tx/{tx_id}/rollback", headers=H) raise r = requests.post(f"{BASE}/tx/{tx_id}/commit", headers=H) if r.status_code == 200: return r.json() if r.status_code == 409 and r.json().get("error") in RETRYABLE: time.sleep(0.05 * (2 ** n)) # back off, then rebuild from scratch continue r.raise_for_status() raise RuntimeError("transaction did not converge after 5 attempts") ``` 7 ## SQL session transactions. `BEGIN`, `COMMIT` and `ROLLBACK` are accepted as statements on `POST /sql`. Write statements issued between them buffer instead of executing, and the whole buffer lands in one frame at `COMMIT`. This is how you make several SQL statements atomic. The buffer is keyed by `(instance, session id)`. Send an explicit `X-OC-Session-Id` header on every statement in the flow. Without it the engine falls back to keying the buffer by your bearer token, which means two concurrent flows sharing a token share a transaction buffer — and neither gets the fence that detects a lost transaction. #### cURL ``` # Every statement in the flow MUST carry the same X-OC-Session-Id. SID=$(uuidgen) curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" -d '{"sql":"BEGIN"}' curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{"sql":"UPDATE shop.orders SET status = $1 WHERE id = $2", "params":["shipped","ord-1001"]}' curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" -d '{"sql":"COMMIT"}' ``` #### Python ``` import uuid SID = str(uuid.uuid4()) TXH = {**H, "X-OC-Session-Id": SID} def sql(stmt, params=None): body = {"sql": stmt} if params is not None: body["params"] = params # positional: $1, $2, ... r = requests.post(f"{BASE}/sql", headers=TXH, json=body) r.raise_for_status() return r.json() sql("BEGIN") sql("UPDATE shop.orders SET status = $1 WHERE id = $2", ["shipped", "ord-1001"]) sql("UPDATE shop.orders SET status = $1 WHERE id = $2", ["shipped", "ord-1002"]) print(sql("COMMIT")["ops_committed"]) ``` #### TypeScript ``` const SID = crypto.randomUUID(); const TXH = { ...H, "X-OC-Session-Id": SID, "Content-Type": "application/json" }; async function sql(stmt: string, params?: unknown[]) { const r = await fetch(`${BASE}/sql`, { method: "POST", headers: TXH, body: JSON.stringify(params ? { sql: stmt, params } : { sql: stmt }), }); if (!r.ok) throw new Error(await r.text()); return r.json(); } await sql("BEGIN"); await sql("UPDATE shop.orders SET status = $1 WHERE id = $2", ["shipped", "ord-1001"]); await sql("UPDATE shop.orders SET status = $1 WHERE id = $2", ["shipped", "ord-1002"]); console.log((await sql("COMMIT")).ops_committed); ``` #### Go ``` sid := uuid.NewString() sqlStmt := func(stmt string, params []any) (map[string]any, error) { payload := map[string]any{"sql": stmt} if params != nil { payload["params"] = params // positional: $1, $2, ... } body, _ := json.Marshal(payload) req, _ := http.NewRequestWithContext(ctx, "POST", base+"/sql", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") req.Header.Set("X-OC-Session-Id", sid) // same id for the whole flow resp, err := http.DefaultClient.Do(req) if err != nil { return nil, err } defer resp.Body.Close() var out map[string]any return out, json.NewDecoder(resp.Body).Decode(&out) } sqlStmt("BEGIN", nil) sqlStmt("UPDATE shop.orders SET status = $1 WHERE id = $2", []any{"shipped", "ord-1001"}) sqlStmt("COMMIT", nil) ``` #### cURL ``` { "kind": "tx", "op": "begin", "ops_committed": 0, "session_id": "..." } ``` #### cURL ``` { "kind": "tx", "op": "commit", "ops_committed": 2, "session_id": "..." } ``` read committed — no conflict detection A SQL session transaction gives you atomicity, not isolation. There is no read set and no snapshot, so two concurrent sessions doing read-modify-write on the same row will both commit and one update is lost — no error, no 409. If a lost update would be a bug, use the [`/tx` endpoints](https://originchaindb.com/docs/schemas/transactions#lifecycle) instead. The level is not selectable here: a `BEGIN` carrying an `ISOLATION LEVEL` clause is refused rather than quietly downgraded. a statement does not see the buffer above it This surface does not read its own writes — a deliberate divergence from PostgreSQL. A `SELECT` after a buffered `INSERT` in the same flow will not see the inserted row, and an `UPDATE … WHERE` matches against committed state, not against what you buffered. One consequence worth spelling out: two `INSERT`s of the same primary key inside one flow collapse into a single row at commit — last one wins, with no duplicate-key error. Sequence read-then-write work so the reads happen before `BEGIN`. the buffer is memory, and it can be lost It lives in the serving process. If the engine restarts, or the instance fails over to its standby, the buffer is gone — the next statement in that flow returns `409` with either "transaction not found — it may have been lost to an engine restart" or "transaction lost on failover". Nothing was written. The recovery is the same in both cases: send `ROLLBACK`, then retry the whole transaction. The engine deliberately refuses to forward the statement to the new primary, because there it would execute as a durable standalone write that your later `COMMIT` or `ROLLBACK` could not undo. a forgotten BEGIN expires after five minutes The buffer is swept once it is five minutes old, measured from `BEGIN` — total age, not idle time. That bound is deliberate: without it, one forgotten `BEGIN` would make every later write on that session silently buffer forever. After the sweep the session reads as not-in-a-transaction, so subsequent statements execute normally and a late `COMMIT` honestly reports that there is nothing open. 8 ## Reading inside a transaction. Point reads by primary key use `GET /tx/:tx_id/rows/:schema/:pk`, shown above. For anything broader there is `POST /tx/:tx_id/query` — with two constraints worth knowing before you plan around it. - It takes a plan document, not a SQL string. There is no way to run a SQL `SELECT` inside a `/tx` transaction today. The body is the same internal plan format the `/v1/query` endpoint accepts. - Full-table scans only. Read-your-writes is wired into the scan leaf and nowhere else, so a plan that would read the store through an index, a column scan, a row-count fast path, or a graph traversal is refused with 400 rather than answered with a stale result. Scans wrapped in filter, project, sort, limit, distinct, aggregate, join or set operators are fine — the operators compose over correct inputs. #### cURL ``` # /tx/:id/query takes a PLAN document, not a SQL string. Every node # carries an "op" tag; the leaf here is a full-table scan. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/tx/$TX_ID/query" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{ "op": "filter", "child": { "op": "scan", "schema": "shop.orders" }, "predicate": { "op": "eq", "path": "customer", "value": "cus-77" } }' ``` #### cURL ``` { "rows": [ ... ], "tx_id": "01JTS0...", "_preview": true } ``` A scan-only read is O(table). If your transaction needs an indexed lookup, do the lookup before `begin` and pass the primary keys in — then read those keys through the transaction so they enter the read set. 9 ## Every error you can get. | HTTP | error | When | What to do | | --- | --- | --- | --- | | 409 | tx_conflict | A key this transaction READ changed between begin and commit. | Retry the whole transaction from begin. Nothing was written. | | 409 | serialization_failure | Serializable only. The commit-time check found a read-write dependency cycle (write skew or a phantom). | Retry the whole transaction from begin. Nothing was written. | | 409 | constraint_violation | A foreign key, back-reference, CHECK or UNIQUE constraint failed at commit. | Not retryable as-is. Fix the data, then run a fresh transaction. The handle goes to rolled_back. | | 409 | (no code) transaction not found | SQL session transactions only. The buffer was lost to an engine restart or a failover. | Send ROLLBACK, then retry the whole transaction. | | 413 | tx_buffer_cap_exceeded | More than 4,096 buffered write operations in one transaction. | Commit and open a new transaction, or use the batch row endpoint instead. | | 429 | tx_open_cap_exceeded | More than 64 concurrently active transactions for the instance. | Commit or roll back what is open. Carries Retry-After: 1. | | 501 | serializable_sharded_unsupported | ?isolation=serializable requested on a sharded instance. | Use snapshot isolation, or run on a single-shard instance. | | 501 | serializable_scan_raft_unsupported | A serializable transaction on a quorum-replicated instance issued a scan or range read. | Read by primary key inside the transaction, or use snapshot isolation for scans. | | 501 | (no code) cross-shard table | A transaction touched a table owned by a different shard. | Keep each transaction inside one shard's tables. | | 400 | (no code) scan-only plan | A /tx/:id/query plan reads the store outside a scan leaf. | Re-express as a scan-based plan, or run the read outside the transaction. | | 400 | (no code) params on BEGIN | A params array was sent with BEGIN, COMMIT or ROLLBACK on /sql. | Transaction-control statements take no bind parameters. Send sql only. | | 500 | (no code) tx commit: … | Validation, encoding or storage failure during commit. | Nothing was written; the handle goes to rolled_back. Retry from a fresh begin. | A `404` from any `/tx` route means the id is unknown or belongs to another instance — the two are intentionally identical. A `409` from a row op means the transaction is no longer `active`; check `GET /tx/:tx_id` for which terminal state it reached. 10 ## Limits & gotchas. | Limit | Value | Why it exists | | --- | --- | --- | | Buffered operations per transaction | 4,096 | The buffer is held whole in memory. Breach returns `413`. For bulk loading use the batch row endpoint, which is a different path with different limits. | | Concurrently active transactions | 64 | Per instance. Each open transaction pins its read and write sets. Breach returns `429` with `Retry-After: 1`. | | Reads per transaction | unbounded | The 4,096 cap counts write operations only. The read set has no cap and is held in memory, so a transaction that reads a very large number of keys is a memory risk you own. | | Buffered statements (SQL session) | unbounded | The SQL session buffer has no operation cap at all. Nothing stops you buffering more than fits in memory — keep these transactions small and let the five-minute age limit be your backstop. | | Request body | 8 MiB | Per request, on every transaction route. A row larger than this cannot be buffered; breach returns `413`. | | Idle timeout (`/tx`) | 5 min | A sweeper aborts transactions with no client activity for this long and marks them `expired`. Every operation refreshes the clock, so an active transaction is never swept. | | Total age (SQL session) | 5 min | Measured from `BEGIN`, not idle time. Activity does not extend it — a SQL session transaction cannot run longer than this, full stop. | one shard per transaction On a sharded instance every table a `/tx` transaction touches — including the targets of any foreign key on a written row — must be owned by the same shard. Anything else is refused with `501`. The check runs before the table is looked up, so you get the real reason rather than a misleading `404`. no schema changes inside a transaction Data-definition statements are not part of the `/tx` surface at all, and schema migrations run through their own online-cutover path. Register or migrate a table before the transaction that uses it. vector, full-text and graph writes are not transactional The `/tx` surface covers rows — there is no `/tx/:id/vector`, `/tx/:id/fts` or `/tx/:id/cypher` route. Embeddings written through the vector endpoint and documents indexed through the full-text endpoint are separate write paths and do not join the transaction: they apply immediately and a rollback does not undo them. Sequence the row write inside the transaction and the derived write after a successful commit. geospatial indexes need a repair after an in-transaction delete Deleting a row inside a transaction does not maintain the spatial index for a geospatial column on that table. If you delete rows this way, run the table's geo reindex endpoint afterwards. Non-transactional deletes are unaffected. read-only transactions still pay for commit A transaction that only read still walks its whole read set at commit. If you are only reading and do not need a consistency check across the reads, skip the transaction and query normally — it is strictly cheaper. no savepoints, no nesting Transactions are flat. There is no `SAVEPOINT`, no partial rollback, and opening a transaction inside another is not a thing — you get two independent transactions that can conflict with each other. ## Related. - [Row CRUD](https://originchaindb.com/docs/schemas/row-crud) — the non-transactional row endpoints these buffer against, and the batch path for bulk writes. - [SQL](https://originchaindb.com/docs/schemas/sql) — the full statement surface, including the write statements a SQL session transaction buffers. - [Error reference](https://originchaindb.com/docs/errors) — the shared error envelope every endpoint returns. - [Multi-node](https://originchaindb.com/docs/multi-node) — what sharding and failover change about the guarantees on this page. --- # Schema tutorial - build your first OriginChainDB schema Canonical source: https://originchaindb.com/docs/schemas/tutorial Sitemap last modified: 2026-09-23T08:33:23.000Z tutorial · schemas # Build your first schema Build a product-catalog schema with columns, a primary key, indexes, and relationships. By the end of this page you'll have a real, working schema with columns, an index, a graph relation, a foreign key, and CHECK constraints. ~15 minutes if you read carefully. 1 ## Name the table. Every schema needs three things at the top: a `namespace`, a `table`, and a `primary_key`. The namespace groups related tables; the table is the name you'll query with; the primary key tells OriginChainDB which column uniquely identifies a row. ``` # manifest.toml namespace = "shop" table = "products" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true ``` We declared `id` as a string column and marked it `required = true` so writes that omit it are rejected. The `primary_key = ["id"]` array names which column (or columns) form the key. For now, single-column. 2 ## Add the data columns. A product needs a name, a category, and a price. Notice how we store price: as `price_cents` with type `i64`, not a float. Floats lose precision on multiplication - $0.10 + $0.20 ≠ $0.30 in float arithmetic. Always store money as integer minor units. ``` namespace = "shop" table = "products" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" [[columns]] name = "category" ty = "str" [[columns]] name = "price_cents" ty = "i64" # money in minor units - never f64 ``` 3 ## Add an index for fast filtering. We'll often want "all products in category X". Without an index, OriginChainDB has to scan every row. An `[[indexes]]` block changes that to a direct lookup. ``` namespace = "shop" table = "products" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" [[columns]] name = "category" ty = "str" [[columns]] name = "price_cents" ty = "i64" # Without this, WHERE category = ... has to scan every row. [[indexes]] name = "by_category" columns = ["category"] ``` Rule of thumb: any column you filter on in `WHERE` or use in graph relations should be indexed. Indexes have a write cost, so don't add them speculatively. 4 ## Link products to suppliers. Every product comes from a supplier. We could just add a `supplier_id` column and join on it. But if we declare it as a `[[relations]]` block, OriginChainDB treats it as a graph edge - we can later walk "all products from supplier X" with a single graph call, no JOIN needed. ``` # (everything from step 3, plus:) [[columns]] name = "supplier_id" ty = "str" # Declare supplier_id as a graph edge. The row write creates the edge # automatically; you can later walk it with the /graph endpoint. [[relations]] name = "supplied_by" from_col = "supplier_id" bidirectional = true [relations.target] namespace = "shop" table = "suppliers" pk = "id" ``` `name` is the verb you'll use to walk the edge ("supplied_by"). `from_col` is the column on this table that points at the target. `bidirectional = true` means the reverse edge ("which products does supplier X supply?") is also indexed automatically. 5 ## Refuse rows pointing at nothing. A relation alone doesn't enforce that `supplier_id` actually exists - you could write a product with `supplier_id = "ghost-42"` and it would happily save. A foreign key rejects that write at the door. ``` # (everything from step 4, plus:) [[foreign_keys]] name = "fk_supplier" from_col = "supplier_id" target_schema = "shop.suppliers" target_col = "id" on_delete = "restrict" # no_action | restrict | set_null ``` `on_delete = "restrict"` means the engine refuses to delete a supplier that still has products pointing at it. Other options: `no_action` (same as restrict at write time), `set_null` (clears the column on the orphaned rows). 6 ## Add data-quality checks. CHECK constraints reject writes that violate rules. Negative prices and unknown category strings are the kind of bugs that quietly corrupt a database over months. Stop them at write time. ``` # (everything from step 5, plus:) [[check_constraints]] name = "price_positive" expression = "price_cents > 0" [[check_constraints]] name = "category_in_set" expression = "category IN ('electronics', 'books', 'shoes')" ``` CHECK expressions support `= != < <= > >=`, `AND/OR/NOT`, `IS [NOT] NULL`, and `IN (...)`. They run on every write - keep them simple to avoid write-time overhead. ## The finished schema. Put it all together. This is what you save as `manifest.toml` and send to OriginChainDB. ``` # manifest.toml - the finished schema. namespace = "shop" table = "products" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" [[columns]] name = "category" ty = "str" [[columns]] name = "price_cents" ty = "i64" [[columns]] name = "supplier_id" ty = "str" [[indexes]] name = "by_category" columns = ["category"] [[relations]] name = "supplied_by" from_col = "supplier_id" bidirectional = true [relations.target] namespace = "shop" table = "suppliers" pk = "id" [[foreign_keys]] name = "fk_supplier" from_col = "supplier_id" target_schema = "shop.suppliers" target_col = "id" on_delete = "restrict" [[check_constraints]] name = "price_positive" expression = "price_cents > 0" [[check_constraints]] name = "category_in_set" expression = "category IN ('electronics', 'books', 'shoes')" ``` To register it, see [Schemas → Register a schema](https://originchaindb.com/docs/schemas) on the overview page. The same `POST /v1/tenants/:t/schemas` call works whether your schema is 10 lines or 200. ## What you can do now. - SQL: `SELECT * FROM shop.products WHERE category = 'shoes'` - fast because of the index. - Graph: "all products supplied by sup-44" via `GET /graph/shop.products/neighbors?rel=supplied_by&pk=sup-44`. - Foreign key enforcement: writes with non-existent supplier IDs return `400 fk_violation`. - CHECK enforcement: negative prices and unknown categories return `400 check_violation`. Want vector search or full-text search on top? Those live on their own runtime endpoints, not in the schema. See [Vector reference](https://originchaindb.com/docs/vector) and [Full-text reference](https://originchaindb.com/docs/fts). --- # Declaring a vector collection - OriginChainDB schema reference Canonical source: https://originchaindb.com/docs/schemas/vector Sitemap last modified: 2026-09-23T08:33:23.000Z schema · vector # Vector schema. Configure vector dimensions, distance, and index settings for embedding search. Reach for it when meaning matters more than wording. If you need exact keyword matching, [full-text search](https://originchaindb.com/docs/schemas/fts) is cheaper and more precise; if you know the predicate, plain [SQL](https://originchaindb.com/docs/schemas/sql) beats both. 1 ## Before you start. Every example on this page uses the `shop.orders` table from the [quickstart](https://originchaindb.com/docs/quickstart), with an embedding of the order's `notes` field attached to each row's id. vectors are not a column type There is no `vector` column type, and vectors cannot be written through the rows endpoint. They live in their own keyspace, addressed by `(tenant, table, id)`, and are written with the dedicated `/vector/…` endpoints [on the Vector reference](https://originchaindb.com/docs/vector#examples). Using the same id on both sides is what lets you take a vector hit and read the full row back with SQL. The `:table` path segment is free-form — it does not have to name a registered schema. Registering one anyway is worth it: it gives you SQL access to the same rows, and the optional `[vector]` block turns a dimension mistake into a readable error. schemas/orders.toml ``` # The row schema. Vectors do NOT live in a column - they sit in their # own keyspace, addressed by the same id. Registering the table is what # lets you read the row back with SQL after a vector hit gives you an id. namespace = "shop" table = "orders" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "customer" ty = "str" [[columns]] name = "amount_cents" ty = "i64" [[columns]] name = "status" ty = "str" [[columns]] name = "notes" ty = "str" [[columns]] name = "placed_ms" ty = "u64" # Optional. Declares the dimensionality the collection expects so a # wrong-sized vector is refused with a readable message instead of a # generic mismatch. `distance` accepts cosine | l2 | dot | manhattan | l1, # defaults to cosine, and resolves the metric queries actually run under. [vector] dim = 768 distance = "cosine" ``` the [vector] block checks dim and metric `dim` is enforced on every write and query. `distance` is enforced too: it resolves the metric a put and a top-k actually run under. Declare a non-default distance (`l2`, `dot`, `manhattan`, `l1`) and send a contradicting `metric`, and both paths return a `400`; declare one and send no `metric`, and the request runs under the declared metric rather than defaulting to cosine. The block accepts all five values — `cosine`, `l2`, `dot`, `manhattan` and `l1` — the same set the runtime `metric` field takes. The one exception is a declared `"cosine"`: it is byte-identical on disk to declaring nothing, so it resolves an omitted metric but never refuses one. 2 ## Metrics and dimensions. | metric | When to use it | | --- | --- | | "cosine" | The default, and the right answer for almost every text embedding model. Angle only — magnitude is ignored. A zero-magnitude vector scores 0. | | "dot" | Inner product. Use it on unit-normalised vectors, where it is equal to cosine. On a corpus with a real magnitude spread it is a known hazard: the graph's neighbour lists end up ordered by vector length, and the smallest-norm vectors lose their last in-edge and stop being reachable. Normalise first, or use cosine. | | "l2" | Euclidean distance, returned negated. Common for image and audio embeddings. Note the accepted string is `l2` — `"euclidean"` is not recognised. | | "manhattan" | L1 distance, returned negated. `"l1"` is accepted as an alias — both spellings are also valid in the `[vector]` schema block. | an unrecognised metric is a 400, not a fallback Metric parsing is case-insensitive, but an unrecognised value is refused: `"euclidean"`, `"hamming"` or a typo like `"cosin"` is a 400 listing the accepted spellings. Until 2026-07-29 all three fell back to cosine behind a `200`, so an older client may still be sending one. An omitted metric is not an error - it resolves to cosine, or to the declared `[vector].distance` where there is one. ### Dimensions. The only hard rule is `dim > 0` — there is no upper bound on dimensionality in the engine. What matters is consistency: every vector in a collection, and every query against it, must agree. A mismatch is a `400`. With a `[vector]` block registered you get the friendly form: ``` vector has 512 dims but collection "shop.orders" expects 768. Use the same model, or re-index the collection. ``` Without one the engine falls back to the width already stored, and the same message reads `already holds vectors of 768` in place of `expects 768`. A different error, `vec: dimension mismatch: expected 768, got 512`, means the request contradicted itself - the `dim` field did not match the length of the array you sent. ### Picking a dimensionality. - Storage is `dim × 4` bytes per raw vector, so 1536-dim costs exactly twice 768-dim before any index overhead. - Graph search cost scales with `dim` too — every distance computation touches every component. - Many modern embedding models support truncation (Matryoshka-style). Halving dimensions usually costs a little recall and saves a lot of memory; measure on your own data before committing, because the collection's `dim` is fixed once you have written vectors at it. ## Related. - [Quickstart — your first top-k](https://originchaindb.com/docs/quickstart#vector) in all four languages. - [Vector examples](https://originchaindb.com/docs/examples/vector) — one focused page per operation. - [Full-text search](https://originchaindb.com/docs/schemas/fts) — the keyword half of a hybrid retrieval stack. - [RAG patterns](https://originchaindb.com/docs/rag) — putting retrieval and generation together. - [Full schema reference](https://originchaindb.com/docs/schemas/reference) — every block the TOML grammar accepts. --- # OriginChainDB SDKs - Python, TypeScript and Go clients Canonical source: https://originchaindb.com/docs/sdk Sitemap last modified: 2026-09-13T20:18:50.000Z sdks # SDKs Three first-party clients over the same HTTP API. Python has the widest surface and is the only one that retries transient failures for you; TypeScript adds a separate control-plane client for instances, access and billing; Go has zero third-party dependencies and wraps the query surfaces only. [Python The widest surface of the three. Sync and async clients, typed namespaces, and the graph algorithms. Retries transient failures for you.](https://originchaindb.com/docs/sdk/python)[TypeScript Data-plane client plus a separate control-plane client for instances, access and billing. ESM and CJS, types included.](https://originchaindb.com/docs/sdk/typescript)[Go Zero third-party dependencies, context on every call, goroutine-safe. The query surfaces, without the schema and row helpers.](https://originchaindb.com/docs/sdk/go) ## What each one covers Every client wraps the same endpoints, but they were not built at the same time and they do not cover the same ground. Read down the column you plan to use, not across. | | Python 0.6.0 | TypeScript 0.4.0 | Go v2.0.0 | | --- | --- | --- | --- | | Schemas: list, get, register | yes | yes | not yet | | Rows: get, put, batch put | yes | not yet | not yet | | SQL: query and single-row | yes | yes | yes | | SQL: execute, materialized views | yes | not yet | not yet | | Vector: put, top-k | yes | yes | yes | | Vector: delete | yes | yes | not yet | | Vector: centroids, training, rebalance | yes | not yet | not yet | | Full-text: index, search | yes | yes | yes | | Full-text: synonyms, stopwords | yes | not yet | not yet | | Graph: neighbours, BFS, path, Dijkstra | yes | yes | yes | | Graph: PageRank, Louvain, betweenness and friends | yes | not yet | not yet | | Natural language ask | yes | yes | yes | | Usage | yes | yes | yes | | Control plane: instances, access, billing | not yet | yes | not yet | | Async variant | yes, AsyncOriginChain | native | context | A gap in a column is not a gap in the product. Every one of those calls is an ordinary HTTP request, documented in [the HTTP API reference](https://originchaindb.com/docs/api) - the client simply has not wrapped it yet. Reach for the endpoint directly and nothing else changes. ## None of them speak Elasticsearch use the elasticsearch client OriginChainDB exposes an Elasticsearch-compatible API, and the point of it is that you keep the client you already have. None of our SDKs wrap it, deliberately - a second hand-rolled search client would drift from the compatibility surface it is meant to mirror. Point the official Elasticsearch client at your instance instead: [the Elasticsearch quickstart](https://originchaindb.com/docs/elasticsearch/quickstart) shows the connection. Elasticsearch is a trademark of Elasticsearch B.V., which is not affiliated with OriginChainDB and does not endorse it. ## Authentication is the same everywhere Three values, whichever client you pick: the endpoint of your instance, a bearer token, and a tenant id. The TypeScript and Go clients derive the tenant from the endpoint's first DNS label, so you usually pass two. Find all three in the console on your instance page. ``` export OC_BASE_URL='https://t-abc.your-region.db.originchain.ai' export OC_BEARER='oc_live_...' export OC_TENANT='t-abc' ``` Tokens are per instance and carry whatever role you issued them with, so [roles and permissions](https://originchaindb.com/docs/dashboard/rbac) apply to SDK calls exactly as they do to raw HTTP. Keep them in a secrets manager; none of the clients read a config file from disk. ## Retries: only one of the three does it for you This is the difference most likely to bite you, so it is worth being blunt about. All three attach an Idempotency-Key to every mutating request, which is what makes a retry safe. Only Python actually performs the retry. | | Automatic idempotency key | Retries transient failures | | --- | --- | --- | | Python | yes | yes - 429, 500, 502, 503 and 504, up to max_retries | | TypeScript | yes | no - write your own loop | | Go | yes | no - write your own loop | In TypeScript and Go, a retry you write yourself is still safe: the key travels with the request, and the engine's server-side cache collapses the duplicate. What you do not get is the loop. ## Packages and source | Language | Install | Source | | --- | --- | --- | | Python | pip install originchain | [originchain-ai/originchain-python](https://github.com/originchain-ai/originchain-python) | | TypeScript | npm install @originchain/sdk | [originchain-ai/originchain-typescript](https://github.com/originchain-ai/originchain-typescript) | | Go | go get github.com/originchain-ai/originchain-go | [originchain-ai/originchain-go](https://github.com/originchain-ai/originchain-go) | Prefer no dependency at all? [The HTTP API](https://originchaindb.com/docs/api) is the whole product: every one of these clients is a wrapper over it, and anything they have not wrapped is one request away. --- # OriginChainDB Go SDK - install, connect and query Canonical source: https://originchaindb.com/docs/sdk/go Sitemap last modified: 2026-09-13T20:18:50.000Z sdks · go # Go SDK No third-party dependencies, a context on every call, and a client that is safe to share across goroutines. Covers the query surfaces; the schema and row helpers are not wrapped yet. ## Install ``` go get github.com/originchain-ai/originchain-go ``` The module depends only on the standard library. ## Connect Build one client and share it. It is safe for concurrent use, reuses connections through the underlying transport, and has no Close method because it holds nothing exclusive. The default HTTP client times out after 30 seconds; supply your own to change that. ``` import ( "context" "os" "github.com/originchain-ai/originchain-go" ) ctx := context.Background() db := originchain.NewClient(originchain.Config{ BaseURL: os.Getenv("OC_BASE_URL"), Bearer: os.Getenv("OC_BEARER"), }) ``` | Field | Meaning | | --- | --- | | BaseURL | engine endpoint - required, and the constructor panics without it | | Bearer | bearer token | | Tenant | optional - derived from the endpoint hostname when empty | | HTTP | supply your own client; the default has a 30 second timeout | the panic is deliberate NewClient panics on an empty BaseURL rather than returning an error, because a client with no endpoint can never do useful work and the mistake is always a wiring bug at startup, not a runtime condition. ## Query ### SQL ``` resp, err := db.SQL(ctx, "SELECT name, price FROM shop.products WHERE price > $1", 500) if err != nil { return err } for _, row := range resp.Rows { fmt.Println(row["name"], row["price"]) } row, err := db.SQLOne(ctx, "SELECT count(*) AS n FROM shop.products") ``` ### Full-text ``` err := db.FTSIndex(ctx, "shop.products", "name", originchain.FTSIndexRequest{ PK: "sku-1", Text: "Aeron chair", }) hits, err := db.FTSSearch(ctx, "shop.products", "name", originchain.FTSSearchRequest{ Q: "chair", Limit: 10, }) ``` ### Vector ``` err := db.VectorPut(ctx, "shop.products", originchain.VectorPutRequest{ PK: "sku-1", Vector: embedding, }) near, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Vector: query, K: 10, }) ``` ### Graph Traversals hang off db.Graph(): ``` n, err := db.Graph().Neighbors(ctx, "shop", originchain.NeighborsRequest{ Rel: "bought_with", PK: "sku-1", }) p, err := db.Graph().Dijkstra(ctx, "shop", originchain.DijkstraRequest{ Rel: "ships_to", From: "a", To: "b", }) ``` ### Ask and usage ``` ans, err := db.Ask(ctx, "which products sold best last week?") u, err := db.Usage(ctx) ``` ## What this client does not wrap yet Schema registration, the row helpers, vector delete, plan queries and the health check have no Go method today. They are ordinary HTTP calls - build them against [the HTTP API reference](https://originchaindb.com/docs/api) with the same bearer, or use the Python client where the helper already exists. ## Errors Errors come back typed. Unwrap them the usual way: ``` var apiErr *originchain.APIError if _, err := db.SQL(ctx, "SELECT 1"); errors.As(err, &apiErr) { if apiErr.Status == 429 { backOff() } } ``` | Type | Raised when | | --- | --- | | *APIError | any non-2xx response - carries the status and the parsed body | | *AddonRequiredError | the call needs an add-on the account does not have | ## Retries and idempotency Every mutating method attaches a fresh UUIDv4 Idempotency-Key generated from crypto/rand, and the engine caches the result server-side. This client does not retry; your own retry policy takes over. That is the safe arrangement: your loop can re-issue the same call and the write still lands once. Set the header yourself when the retry has to survive a restart. ## Source go get github.com/originchain-ai/originchain-go · [originchain-ai/originchain-go](https://github.com/originchain-ai/originchain-go) · [the HTTP API underneath](https://originchaindb.com/docs/api) --- # OriginChainDB Python SDK - install, connect and query Canonical source: https://originchaindb.com/docs/sdk/python Sitemap last modified: 2026-09-13T20:18:50.000Z sdks · python # Python SDK The widest of the three clients. Sync and async, typed namespaces for every query surface, and the only one that retries transient failures for you. ## Install Python 3.9 or newer. The client uses httpx, and will negotiate HTTP/2 when the extra is present. ``` pip install originchain ``` ## Connect The standard bootstrap reads three environment variables - OC_BASE_URL, OC_BEARER and OC_TENANT - and raises if any is missing, rather than failing later on the first call. ``` from originchain import OriginChain db = OriginChain.from_env() print(db.health()) ``` Pass them explicitly when the environment is not yours to set, in tests for instance: ``` db = OriginChain( base_url="https://t-abc.your-region.db.originchain.ai", bearer="oc_live_...", tenant="t-abc", timeout=30.0, max_retries=3, ) ``` There is an async client with the same surface. Everything on this page works on it with await, and it closes with await db.aclose(). ``` from originchain import AsyncOriginChain db = AsyncOriginChain.from_env() rows = await db.sql("SELECT 1") await db.aclose() ``` ## The namespaces The client is organised by surface. Each namespace hangs off the client instance: | Namespace | What is on it | | --- | --- | | db.schemas | list(), get(), register() | | db.rows | get(), put(), put_batch() | | db.sql | callable, plus query(), execute() and the materialized-view calls | | db.vector | put(), topk(), delete(), delete_bulk(), centroid install and training | | db.fts | index(), search(), install_synonyms(), install_stopwords() | | db.graph | neighbours, BFS, shortest path, Dijkstra, and the graph algorithms | | db.admin | per-tenant replication configuration | one attribute, two shapes db.sql is a callable namespace. db.sql("SELECT ...") and db.sql.query("SELECT ...") both work and hit the same endpoint, so older code keeps running while new code can reach the typed calls next to it. ## Register a schema and write rows A schema is TOML, and it is the same declaration every surface reads - see [the schema reference](https://originchaindb.com/docs/schemas/reference). ``` db.schemas.register(open("products.toml").read()) db.schemas.list() db.rows.put("shop.products", { "id": "sku-1", "name": "Aeron chair", "price": 1195, }) # Batched writes are chunked, and each chunk carries its own # derived idempotency key, so a partial retry does not duplicate. db.rows.put_batch("shop.products", rows) ``` ## Query ### SQL ``` resp = db.sql.query( "SELECT name, price FROM shop.products WHERE price > 500 ORDER BY price DESC" ) for r in resp.rows: print(r["name"], r["price"]) one = db.sql_one("SELECT count(*) AS n FROM shop.products") ``` ### Full-text ``` db.fts.index("shop.products", "name", pk="sku-1", text="Aeron chair") hits = db.fts.search("shop.products", "name", q="chair", limit=10) ``` ### Vector ``` db.vector.put("shop.products", pk="sku-1", vector=embedding) hits = db.vector.topk("shop.products", vector=query_vec, k=10) ``` ### Graph The five traversals plus the algorithm calls. Which algorithms your instance answers is covered in [the graph reference](https://originchaindb.com/docs/graph). ``` db.graph.neighbors("shop", rel="bought_with", pk="sku-1") db.graph.bfs("shop", rel="bought_with", pk="sku-1", depth=3) db.graph.pagerank("shop", rel="bought_with") db.graph.louvain("shop", rel="bought_with") ``` ### Natural language ``` answer = db.ask("which products sold best last week?", schemas=["shop"]) ``` ## Errors Every failure is a typed exception, so you can catch the one case you can do something about: | Exception | Raised when | | --- | --- | | OCAuthError | the token is missing, wrong, or lacks the role for this call | | OCValidationError | the request was malformed - a bad schema, an unknown column | | OCNotFoundError | the schema, table or row does not exist | | OCPaymentRequiredError | the call needs an add-on the account does not have | | OCRateLimitedError | the tenant is over its request budget | | OCServerError | the engine failed - retried already, if retries are on | | OCReplicationDegraded | a warning, not an exception: the read was served while a replica was catching up | All of them subclass OCError, so a single catch still works. ``` from originchain import OCRateLimitedError, OCError try: db.sql("SELECT 1") except OCRateLimitedError: back_off() except OCError as e: log.exception(e) ``` ## Retries and idempotency This client retries 429, 500, 502, 503 and 504 up to max_retries times. Mutating requests carry a generated Idempotency-Key, which the engine caches server-side, so a retried write lands once. Pass a stable idempotency_key yourself when the retry has to survive a process restart - a job runner picking the same task up again, for instance. Batched writes derive one key per chunk from the key you give, so partial progress is preserved. ## Source pip install originchain · [originchain-ai/originchain-python](https://github.com/originchain-ai/originchain-python) · [the HTTP API underneath](https://originchaindb.com/docs/api) --- # OriginChainDB TypeScript SDK - install, connect and query Canonical source: https://originchaindb.com/docs/sdk/typescript Sitemap last modified: 2026-09-13T20:18:50.000Z sdks · typescript # TypeScript SDK A typed client over the query surfaces, plus a second client for the control plane. Ships ESM and CommonJS builds with declarations, and runs anywhere there is a fetch. ## Install Node 18 or newer, where fetch is built in. Works with npm, pnpm, yarn and bun. ``` npm install @originchain/sdk ``` ## Connect Two required options. The tenant id is parsed from the first DNS label of the endpoint, so you only pass tenantId for non-standard hostnames or local development. Every request times out after 60 seconds unless you set timeoutMs. ``` import { OriginChainClient } from "@originchain/sdk"; const db = new OriginChainClient({ baseUrl: process.env.OC_BASE_URL!, bearer: process.env.OC_BEARER!, }); ``` | Option | Meaning | | --- | --- | | baseUrl | engine endpoint - required, throws if absent | | bearer | bearer token - required, throws if absent | | tenantId | override the tenant derived from the hostname | | timeoutMs | per-request timeout, 60 seconds by default | | fetch | swap the fetch implementation, for tests or instrumentation | ## Query ### SQL ``` const resp = await db.sql( "SELECT name, price FROM shop.products WHERE price > $1", [500] ); const one = await db.sqlOne("SELECT count(*) AS n FROM shop.products"); ``` ### Full-text ``` await db.ftsIndex("shop.products", "name", { pk: "sku-1", text: "Aeron chair" }); const hits = await db.ftsSearch("shop.products", "name", { q: "chair", limit: 10 }); ``` ### Vector ``` await db.vectorPut("shop.products", { pk: "sku-1", vector: embedding }); const near = await db.vectorTopk("shop.products", { vector: q, k: 10 }); await db.vectorDelete("shop.products", { pk: "sku-1" }); ``` ### Graph The traversals live on a graph property, not on the client itself: ``` await db.graph.neighbors("shop", { rel: "bought_with", pk: "sku-1" }); await db.graph.bfs("shop", { rel: "bought_with", pk: "sku-1", depth: 3 }); await db.graph.dijkstra("shop", { rel: "ships_to", from: "a", to: "b" }); ``` ### Schemas, ask and usage ``` await db.listSchemas(); await db.registerSchema(toml); await db.ask("which products sold best last week?"); await db.usage(); ``` no row helpers here This client has no rows.put equivalent yet - write rows through SQL, or call the row endpoint directly from [the HTTP API](https://originchaindb.com/docs/api). The Python client has the helper if you need it today. ## The control-plane client Instances, access, billing and operations live behind a different service and a different client. It is exported from the same package. ``` import { OriginChainAdminClient } from "@originchain/sdk"; const cp = new OriginChainAdminClient({ baseUrl: "https://api.originchain.ai", credentials: "omit", bearer: token, }); await cp.instances.list(); await cp.instances.setAllowlist(id, entries); await cp.instances.enablePgwire(id); await cp.instances.metrics(id, 60, 60); ``` It also covers sign-in and session management, snapshots, point-in-time archives, logs and per-instance schemas. Cookies are sent by default for browser use - pass credentials: "omit" and an explicit bearer from Node. ## Errors Failures throw. The base class carries the status and the parsed body: | Class | Raised when | | --- | --- | | ApiError | any non-2xx response - has the status and body on it | | OCAddonRequiredError | the call needs an add-on the account does not have | | OCPaymentRequiredError | payment is required before the call can proceed | ``` import { ApiError } from "@originchain/sdk"; try { await db.sql("SELECT 1"); } catch (e) { if (e instanceof ApiError && e.status === 429) await backOff(); else throw e; } ``` ## Retries and idempotency Every mutating request gets a generated Idempotency-Key header, and the engine caches the result server-side. This client does not retry for you - it applies the per-request timeout and throws. So write the loop yourself. Because the key travels with the request, a retry of the same call collapses into the original write rather than duplicating it. Set the header explicitly when a retry must survive a process restart. ## Source npm install @originchain/sdk · [originchain-ai/originchain-typescript](https://github.com/originchain-ai/originchain-typescript) · [the HTTP API underneath](https://originchaindb.com/docs/api) --- # SQL on OriginChainDB - SELECT, GROUP BY, JOIN Canonical source: https://originchaindb.com/docs/sql Sitemap last modified: 2026-09-23T08:33:23.000Z reference · sql # SQL reference Filter, join, and aggregate rows with OriginChainDB SQL, using the supported syntax and client examples below. All examples below assume you have a client set up. If you haven't yet, see [Quickstart](https://originchaindb.com/docs/quickstart). ## At a glance. works today - SELECT — projection, `*`, `DISTINCT` - WHERE — `=` `!=` `<` `<=` `>` `>=` `BETWEEN` `IN` `IS NULL` `LIKE`, combined with `AND` / `OR` / `NOT` - ORDER BY (multi-column, ASC/DESC), `LIMIT`, `OFFSET` - GROUP BY + COUNT, SUM, AVG, MIN, MAX, `HAVING`, `COUNT(DISTINCT)` - JOIN — INNER, LEFT, RIGHT, FULL OUTER (up to 32 tables), and JOIN + GROUP BY - Window functions — ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD, SUM/AVG `OVER (PARTITION BY)`, with `ROWS` / `RANGE` frame clauses - Subqueries — `IN` / `EXISTS` / scalar, correlated and uncorrelated (across tables) - CTEs — `WITH … AS` over real tables, with filters, joins and aggregates, plus a single `WITH RECURSIVE` (`UNION ALL` only) - Set operations — `UNION`, `UNION ALL`, `INTERSECT`, `EXCEPT`, chainable across three or more SELECTs - CASE expressions — `CASE WHEN … THEN … ELSE … END` in the SELECT list - INSERT, UPDATE, DELETE — set-based, matched by a `WHERE` across many rows - Transactions — BEGIN / COMMIT / ROLLBACK, atomic even across nodes — see [Transactions](https://originchaindb.com/docs/transactions) - DDL — `CREATE` / `DROP` TABLE & SCHEMA (`IF NOT EXISTS`), `PRIMARY KEY`, `DEFAULT`, `FOREIGN KEY` (inline + table), `SERIAL` / `IDENTITY`, `ALTER … ADD CONSTRAINT`, `CREATE INDEX` / `VIEW`, `EXPLAIN` - Parameters — positional `$1` bound from a `params` array - Catalog & session — `pg_catalog` / `information_schema` introspection, `SET` / `SHOW` (search_path, GUCs), `current_schema()` - Time-bucketing — `date_trunc` over a timestamp column in `GROUP BY` (the dashboard shape) not yet - Scalar subquery in the SELECT list — `SELECT (SELECT …) AS x`; in a `WHERE` comparison it ships - Qualified projection off a derived table — `SELECT d.c FROM (…) d`; `SELECT c` and `SELECT *` over the same derived table work - nextval() as a column `DEFAULT` or `INSERT` value — supply an epoch literal or a bound param Try anything in the "not yet" list and you get a `400` with a clear reason. Nothing is silently re-interpreted. ## The endpoint. ``` POST /v1/tenants/:tenant/sql Authorization: Bearer $OC_TOKEN Content-Type: application/json { "sql": "SELECT ..." } ``` The response shape depends on the statement kind. Every response carries a `"kind"` field so your code can switch on it: | kind | When | What's in the body | | --- | --- | --- | | "select" | SELECT | `rows: [{...}, ...]` - one object per row. | | "explain" | EXPLAIN | `plan` - the pretty-printed plan tree, plus `stats` on `EXPLAIN ANALYZE`. EXPLAIN does not carry kind "select". | | "insert" | INSERT | `schema`, `rows` - the inserted rows. INSERT executes and enforces foreign keys. See [Writes](https://originchaindb.com/docs/sql#writes) below. | | "update" | UPDATE | `rows_affected` - UPDATE executes for every row the WHERE clause matches. A predicate on the primary key takes a fast path; any other supported predicate lowers to a scan. | | "delete" | DELETE | `schema`, `pk`, `rows_affected` - DELETE executes for every row the WHERE clause matches and is durable when the response returns; a transaction is not required. See [Writes](https://originchaindb.com/docs/sql#writes). | `` `` ``Column types. You declare a column's type when you register a schema (see schemas) or with CREATE TABLE. The engine validates every value against it. These are the types it stores: Type SQL aliases Holds i64 / u64INT, BIGINT, SMALLINTSigned / unsigned integers. f64FLOAT, REAL, DOUBLEFloating-point numbers. decimalDECIMAL(p,s), NUMERICExact fixed-point (money). str / textTEXT, VARCHAR, CHARUTF-8 strings. boolBOOL, BOOLEANTrue / false. timestamp / dateTIMESTAMP, DATETIME, DATEInstants / calendar dates. uuidUUIDFormat-validated UUIDs. enumENUM(...)One of a fixed set of strings. jsonJSON, JSONBArbitrary JSON values. listARRAYOrdered lists. bytesBLOB, BYTEA, BINARYRaw byte strings. inet / interval / pointINET, INTERVAL, POINTIP addresses, durations, geo points. on this page SELECT basics WHERE filters GROUP BY + aggregates JOIN Writes via SQL EXPLAIN The endpoint contract SELECT Bind parameters Aggregates & GROUP BY Joins Writes and schema changes EXPLAIN Where GROUP BY, JOIN and HAVING run Limits & gotchas 1. SELECT basics. what this does Read rows back from a table. Pick which columns you want, filter with WHERE, cap the result count with LIMIT. cURL Python TypeScript Go POST /v1/tenants/:t/sql curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, name, price_cents FROM shop.products WHERE category = '\''electronics'\'' LIMIT 50" }'result = db.sql(""" SELECT id, name, price_cents FROM shop.products WHERE category = 'electronics' LIMIT 50 """) # result is a SqlSelect when the statement was a SELECT. for row in result.rows: print(row["id"], row["name"], row["price_cents"])const result = await db.sql(` SELECT id, name, price_cents FROM shop.products WHERE category = 'electronics' LIMIT 50 `); if (result.kind === "select") { for (const row of result.rows) { console.log(row.id, row.name, row.price_cents); } }result, err := db.SQL(ctx, ` SELECT id, name, price_cents FROM shop.products WHERE category = 'electronics' LIMIT 50 `) if err != nil { /* handle */ } if result.Kind == "select" { for _, row := range result.Rows { fmt.Println(row["id"], row["name"], row["price_cents"]) } } common mistakes ORDER BY runs server-side. Sort by one or more columns with ASC/DESC, and page through results with LIMIT + OFFSET. (Window frame clauses like ROWS BETWEEN are supported; GROUPS mode and EXCLUDE are not.) Missing LIMIT. A SELECT without LIMIT returns every matching row. For large tables this can be slow. Always include a LIMIT during development. Schema vs. table name. Use the full schema.table form (here, shop.products). The bare table name without the schema doesn't resolve. 2. WHERE filters. what this does Narrow down which rows come back. Combine any number of conditions with AND, OR and NOT. cURL Python TypeScript Go POST /v1/tenants/:t/sql curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, name, price_cents FROM shop.products WHERE category = '\''electronics'\'' AND price_cents > 10000 AND price_cents < 50000" }'# Multiple AND conditions. WHERE supports = != < <= > >=, BETWEEN, # IN, IS NULL, IS NOT NULL, and LIKE. Combine with AND, OR and NOT. result = db.sql(""" SELECT id, name, price_cents FROM shop.products WHERE category = 'electronics' AND price_cents > 10000 AND price_cents < 50000 """)// Multiple AND conditions. WHERE supports = != < <= > >=, BETWEEN, // IN, IS NULL, IS NOT NULL, and LIKE. Combine with AND, OR and NOT. const result = await db.sql(` SELECT id, name, price_cents FROM shop.products WHERE category = 'electronics' AND price_cents > 10000 AND price_cents < 50000 `);// Multiple AND conditions. WHERE supports = != < <= > >=, BETWEEN, // IN, IS NULL, IS NOT NULL, and LIKE. Combine with AND, OR and NOT. result, _ := db.SQL(ctx, ` SELECT id, name, price_cents FROM shop.products WHERE category = 'electronics' AND price_cents > 10000 AND price_cents < 50000 `) operators you can use Operator Example = != < <= > >= price_cents > 1000 BETWEEN price_cents BETWEEN 1000 AND 20000 IN (list) category IN ('electronics', 'books') IS NULL / IS NOT NULL description IS NOT NULL LIKE name LIKE 'Wireless%' common mistakes AND, OR, and NOT all work. Combine them freely — e.g. WHERE (category = 'books' OR category = 'toys') AND price_cents < 5000. IN (...) is still the tidiest way to match a set of values. No bind parameters. Inline literal values - $1 and ? placeholders are not supported yet. If you build SQL from user input, escape strings carefully. String quoting in cURL. Single quotes inside a JSON string need to be escaped as '\\''. Easier to use a Python / TS / Go SDK for any non-trivial query. 3. GROUP BY + aggregates. what this does Roll rows up by one or more columns and compute aggregates (counts, sums, averages, min/max). Useful for any "how many X per Y" question. cURL Python TypeScript Go POST /v1/tenants/:t/sql curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT category, COUNT(*) AS n, SUM(price_cents) AS total, AVG(price_cents) AS avg_price FROM shop.products GROUP BY category" }'result = db.sql(""" SELECT category, COUNT(*) AS n, SUM(price_cents) AS total, AVG(price_cents) AS avg_price FROM shop.products GROUP BY category """) for row in result.rows: print(row["category"], row["n"], row["total"], row["avg_price"])const result = await db.sql(` SELECT category, COUNT(*) AS n, SUM(price_cents) AS total, AVG(price_cents) AS avg_price FROM shop.products GROUP BY category `); if (result.kind === "select") { for (const row of result.rows) { console.log(row.category, row.n, row.total, row.avg_price); } }result, _ := db.SQL(ctx, ` SELECT category, COUNT(*) AS n, SUM(price_cents) AS total, AVG(price_cents) AS avg_price FROM shop.products GROUP BY category `) if result.Kind == "select" { for _, row := range result.Rows { fmt.Println(row["category"], row["n"], row["total"], row["avg_price"]) } } supported aggregate functions Function What it returns COUNT(*) Number of rows in the group. COUNT(col) Number of non-null values of col. SUM(col) Sum of values. AVG(col) Arithmetic mean of values. MIN(col), MAX(col) Smallest / largest value. common mistakes Use HAVING to filter groups. WHERE filters rows before grouping; HAVING filters groups after — e.g. SELECT category, COUNT(*) AS n … GROUP BY category HAVING COUNT(*) > 3. The aggregate you filter on has to appear in the SELECT list too. Both run server-side. DISTINCT aggregates work. COUNT(DISTINCT col), SUM(DISTINCT col), and AVG(DISTINCT col) all execute. (COUNT(DISTINCT *) is the one form that isn't allowed.) Every selected column needs to be in GROUP BY or an aggregate. Standard SQL rule - if you select a column you didn't group by, you'll see 400. 4. JOIN tables. what this does Combine rows from two or more tables on a matching column. INNER, LEFT, RIGHT, and FULL OUTER joins are supported. Up to 32 tables in one query. cURL Python TypeScript Go POST /v1/tenants/:t/sql curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT o.id, o.qty, p.name FROM shop.orders o INNER JOIN shop.products p ON o.product_id = p.id WHERE o.status = '\''paid'\'' LIMIT 100" }'result = db.sql(""" SELECT o.id, o.qty, p.name FROM shop.orders o INNER JOIN shop.products p ON o.product_id = p.id WHERE o.status = 'paid' LIMIT 100 """) for row in result.rows: print(row["o.id"], row["o.qty"], row["p.name"])const result = await db.sql(` SELECT o.id, o.qty, p.name FROM shop.orders o INNER JOIN shop.products p ON o.product_id = p.id WHERE o.status = 'paid' LIMIT 100 `); if (result.kind === "select") { for (const row of result.rows) { console.log(row["o.id"], row["o.qty"], row["p.name"]); } }result, _ := db.SQL(ctx, ` SELECT o.id, o.qty, p.name FROM shop.orders o INNER JOIN shop.products p ON o.product_id = p.id WHERE o.status = 'paid' LIMIT 100 `) if result.Kind == "select" { for _, row := range result.Rows { fmt.Println(row["o.id"], row["o.qty"], row["p.name"]) } } join types Type What you get INNER JOIN Only rows that have a match on both sides. LEFT JOIN Every row from the left side, plus matches from the right (null if no match). RIGHT JOIN Mirror of LEFT - every row from the right side, plus matches from the left. FULL OUTER JOIN Every row from both sides; nulls fill the gaps. common mistakes CROSS JOIN is narrow. Only the two-table SELECT * FROM a CROSS JOIN b form works, optionally with a LIMIT. A third table, a mix with INNER / OUTER joins, or a WHERE, ORDER BY, OFFSET, DISTINCT or explicit projection on one returns 400 — use an explicit ON condition instead. Ambiguous column names. When two tables have the same column name, qualify with the alias (o.id, p.id) - otherwise the parser refuses. 33+ tables. The cap is 32 tables per query. Larger joins return 400. (You almost never want more than 5 in practice.) 5. Writes via SQL (preview). INSERT and UPDATE execute against the engine. INSERT writes the rows and enforces foreign keys; UPDATE changes every row its WHERE clause matches, taking a fast path when that predicate is the primary key and returns rows_affected. For high-volume ingest, the dedicated row endpoints are still the fastest path. cURL POST /v1/tenants/:t/sql - INSERT curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "INSERT INTO shop.products (id, name, category, price_cents) VALUES ('\''p-x-1'\'', '\''New Product'\'', '\''electronics'\'', 5555)" }' response { "kind": "insert", "schema": "shop.products", "rows": [ { "id": "p-x-1", "name": "New Product", "category": "electronics", "price_cents": 5555 } ] } deletes You can run row deletes inside a transaction — BEGIN; DELETE FROM shop.products WHERE id = 'p-x-1'; COMMIT; — which buffers the delete and applies it on commit. A bare DELETE outside a transaction executes immediately and is durable when the response returns. Both UPDATE and DELETE are set-based — they act on every row the WHERE clause matches. See Transactions for the full lifecycle, error handling, and retry pattern. 6. EXPLAIN. what this does Prefix any SELECT with EXPLAIN to see the query plan the engine would run, without executing it. Useful for checking whether your indexes are being used. cURL Python TypeScript Go POST /v1/tenants/:t/sql curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "EXPLAIN SELECT id, name FROM shop.products WHERE category = '\''electronics'\''" }'result = db.sql("EXPLAIN SELECT id, name FROM shop.products WHERE category = 'electronics'") print(result.rows[0]) # → the plan tree as JSONconst result = await db.sql( `EXPLAIN SELECT id, name FROM shop.products WHERE category = 'electronics'` ); if (result.kind === "select") console.log(result.rows[0]);result, _ := db.SQL(ctx, "EXPLAIN SELECT id, name FROM shop.products WHERE category = 'electronics'") if result.Kind == "select" { fmt.Println(result.Rows[0]) } The plan tree includes operator names like Scan, Filter, IndexScan, HashJoin, Aggregate. If you see Scan where you expected IndexScan, you likely need an [[indexes]] declaration on the schema. Examples. SQL is the general-purpose surface: reach for it whenever the question is which rows rather than this row — filtering, sorting, counting, summing, joining — and for the writes the row endpoints do not cover, like a partial update or a delete. One endpoint takes a statement, executes it, and returns JSON. It is a large, real SQL subset rather than a full dialect, and this page is written to be exact about the boundary. Everything listed as supported was read off the translator and the executor; everything refused is refused with an error, not silently ignored. 7 The endpoint contract. POST /v1/tenants/:tenant/sql. The body has exactly two fields — sql (required) and params (optional). There is no namespace field, no limit field and no result-format field; the statement carries all of that. One statement per request. Two statements separated by a semicolon are refused. Every success carries a kind discriminator naming what ran — select, insert, update, delete, explain, tx, or one of the schema-change kinds. Branch on it; do not assume a rows array is there. row keys are alphabetical, not projection order Each row is a JSON object, and its keys serialize in alphabetical order — SELECT id, customer comes back with customer first. That is why a columns array is included on the response: it carries the real projection order. Anything that renders a table or writes a CSV should read columns, not the key order of the first row. columns is omitted for SELECT * The array is only present when the plan declares a static projection order. SELECT *, a join wildcard, and a set operation each declare none, so the field is left off the response entirely — not sent as an empty array. If column order matters to you, list the columns explicitly. 8 SELECT. cURL Python TypeScript Go projection with WHERE, ORDER BY and LIMIT curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, customer, amount_cents FROM shop.orders WHERE status = '\''paid'\'' ORDER BY placed_ms DESC LIMIT 3" }'result = db.sql.query(""" SELECT id, customer, amount_cents FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 3 """) print(result.columns) # -> ['id', 'customer', 'amount_cents'] for row in result.rows: print(row["id"], row["amount_cents"])const res = await db.sql(` SELECT id, customer, amount_cents FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 3 `); if (res.kind === "select") { for (const row of res.rows as Record[]) { console.log(row.id, row.amount_cents); } }res, err := db.SQL(ctx, ` SELECT id, customer, amount_cents FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 3 `) if err != nil { /* handle */ } if res.Kind == "select" { for _, row := range res.Rows { fmt.Println(row["id"], row["amount_cents"]) } }response200{ "kind": "select", "columns": ["id", "customer", "amount_cents"], "rows": [ { "amount_cents": 8900, "customer": "cus-12", "id": "ord-2002" }, { "amount_cents": 4200, "customer": "cus-77", "id": "ord-1001" }, { "amount_cents": 1500, "customer": "cus-12", "id": "ord-2001" } ] }Note the keys inside each row are ALPHABETICAL, not projection order. "columns" is the projection order - use it if order matters. Clause by clause Clause Runs Boundaries SELECT * / column list / expressions / AS aliases yes Mixing * with expression projections is refused — list the columns. SELECT DISTINCT yes DISTINCT ON (…) is refused. DISTINCT with GROUP BY or a window function is refused. WHERE yes See the operator table above. ORDER BY col [ASC|DESC], … yes Bare column names or a projection position only. Functions and expressions are refused, and there is no NULLS FIRST / NULLS LAST. LIMIT n [OFFSET m] yes OFFSET without LIMIT works. FETCH FIRST … ROWS ONLY is refused — use LIMIT. GROUP BY … / HAVING … yes Included from the Thunder configuration up. ROLLUP / CUBE / GROUPING SETS are refused. JOIN (INNER / LEFT / RIGHT / FULL / CROSS) yes Included from the Thunder configuration up. See the join rules below. Subqueries in WHERE yes IN, EXISTS and scalar =, correlated or not. One level of nesting. A scalar subquery in the SELECT list is refused. Over a derived table (FROM (SELECT …)) an unqualified SELECT c or SELECT * works; a qualified SELECT d.c is refused. WITH … AS (…) · WITH RECURSIVE yes Recursive form is base UNION ALL recursive, depth-capped at 100. Nested WITH and column-list renaming are refused. UNION · UNION ALL · INTERSECT · EXCEPT yes Outer ORDER BY / LIMIT / OFFSET wrap the combined result. Window functions OVER (…) yes ROW_NUMBER, RANK, DENSE_RANK, LAG, LEAD, FIRST_VALUE, LAST_VALUE, NTILE, and SUM / AVG / MIN / MAX / COUNT. Single-table SELECT list only — refused alongside JOIN, GROUP BY or DISTINCT. CASE WHEN … THEN … END yes Both the searched and the simple form. CAST(x AS t) · x::t yes TRY_CAST and SAFE_CAST are refused. SELECT with no FROM partly A constant expression like SELECT 1 + 2 AS three works. A bare SELECT 1 that a driver sends as a liveness probe does not — every SELECT that names a column must name a table. WHERE operators Operator Notes = != <> < <= > >= Against a literal, another column, or an expression. AND · OR · NOT · ( ) Arbitrary boolean trees. IN (…) · NOT IN (…) Literal lists and subqueries, correlated or not. Three-valued null logic. BETWEEN a AND b Closed interval, and the form that gets index range pushdown. NOT BETWEEN is refused. IS NULL · IS NOT NULL LIKE · NOT LIKE · ILIKE · NOT ILIKE The ESCAPE clause is refused. EXISTS · NOT EXISTS Correlated and uncorrelated. = (SELECT …) Scalar subquery. Only with =; the ordering comparators are refused. Subqueries of every form are refused outright on a tenant with read-side RBAC configured, or an identity-less READ posture of deny - the subquery executor cannot enforce per-user reads, so it fails closed. Scalar functions usable in a projection or a predicate: NOW() / CURRENT_TIMESTAMP · LOWER · UPPER · LENGTH / CHAR_LENGTH · COALESCE · NULLIF · ABS · ROUND · FLOOR · CEIL / CEILING · MOD · POWER / POW · SQRT · CONCAT · SUBSTRING / SUBSTR · TRIM / BTRIM / LTRIM / RTRIM · REPLACE · POSITION / STRPOS Anything outside that list is refused with a message enumerating what is available. Note two absences people reach for: the || string-concatenation operator is not available in a row expression — use CONCAT(a, b) — and there are no date-part extraction functions. 9 Bind parameters. Placeholders are PostgreSQL-style and positional: $1, $2, and so on, filled from the params array in order. Values are substituted into the parsed statement as literals — never spliced into the text — so a parameter can change what a query matches but can never change what the statement does. Use them for anything that came from a user. cURL Python TypeScript Go parameterised SELECT # Placeholders are PostgreSQL-style and positional: $1, $2, ... # params[0] fills $1. JDBC-style "?" is refused. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT id, amount_cents FROM shop.orders WHERE status = $1 AND amount_cents > $2 LIMIT $3", "params": ["paid", 1000, 20] }'import os, requests BASE = f"https://{os.environ['OC_HOST']}/v1/tenants/{os.environ['OC_TENANT']}" H = {"Authorization": f"Bearer {os.environ['OC_TOKEN']}"} # The Python client's params= argument sends a NAMED mapping, which this # engine does not accept - it binds positional $1, $2 from a JSON array. # Until the client is updated, bind through the HTTP API directly: res = requests.post(f"{BASE}/sql", headers=H, json={ "sql": "SELECT id, amount_cents FROM shop.orders " "WHERE status = $1 AND amount_cents > $2 LIMIT $3", "params": ["paid", 1000, 20], }) res.raise_for_status() for row in res.json()["rows"]: print(row["id"], row["amount_cents"])// The TypeScript client takes a positional array and sends it as-is. const res = await db.sql( `SELECT id, amount_cents FROM shop.orders WHERE status = $1 AND amount_cents > $2 LIMIT $3`, ["paid", 1000, 20], ); if (res.kind === "select") console.log(res.rows.length);// The Go client's variadic params are sent as a positional JSON array. // (Its doc comment predates parameter support in the engine.) res, err := db.SQL(ctx, `SELECT id, amount_cents FROM shop.orders WHERE status = $1 AND amount_cents > $2 LIMIT $3`, "paid", 1000, 20, ) if err != nil { /* handle */ } fmt.Println(len(res.Rows)) Scalars only. Strings, numbers, booleans and null. A JSON array or object as a parameter is refused — you cannot bind a list to IN ($1); generate IN ($1, $2, $3) instead. The count must match exactly, both ways. A supplied value the statement never references is an error, and so is a $3 with only two values. Referencing the same placeholder twice is fine. ? is not accepted. The error says so explicitly. Numbering starts at $1 — $0 is out of range. They work in LIMIT too, and in INSERT … VALUES and UPDATE … SET. Not on transaction verbs. Sending params alongside BEGIN, COMMIT or ROLLBACK is a 400. client support is uneven right now The TypeScript and Go clients both send a positional array, which is what the engine binds. The Python client's params= argument sends a named mapping against :name placeholders — a shape this engine does not accept, so the request is rejected before the statement is parsed. From Python, bind through the HTTP API as shown above until the client is updated. 10 Aggregates & GROUP BY. There are exactly five aggregate functions: COUNT, SUM, AVG, MIN and MAX. Anything else — standard deviation, string aggregation, percentiles — is refused with that list in the message. cURL Python TypeScript Go GROUP BY with HAVING curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total FROM shop.orders WHERE status = '\''paid'\'' GROUP BY customer HAVING SUM(amount_cents) > 5000 ORDER BY total DESC LIMIT 10" }'result = db.sql.query(""" SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total FROM shop.orders WHERE status = 'paid' GROUP BY customer HAVING SUM(amount_cents) > 5000 ORDER BY total DESC LIMIT 10 """) for row in result.rows: print(row["customer"], row["orders"], row["total"])const res = await db.sql(` SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total FROM shop.orders WHERE status = 'paid' GROUP BY customer HAVING SUM(amount_cents) > 5000 ORDER BY total DESC LIMIT 10 `); if (res.kind === "select") console.table(res.rows);res, err := db.SQL(ctx, ` SELECT customer, COUNT(*) AS orders, SUM(amount_cents) AS total FROM shop.orders WHERE status = 'paid' GROUP BY customer HAVING SUM(amount_cents) > 5000 ORDER BY total DESC LIMIT 10 `) if err != nil { /* handle */ } for _, row := range res.Rows { fmt.Println(row["customer"], row["orders"], row["total"]) }response200{ "kind": "select", "columns": ["customer", "orders", "total"], "rows": [ { "customer": "cus-12", "orders": 2, "total": 10400 } ] }GROUP BY and HAVING are included from the Thunder configuration up - see section 14. COUNT(*) is the only wildcard form. COUNT(DISTINCT x), SUM(DISTINCT x) and AVG(DISTINCT x) all work; DISTINCT on MIN or MAX is accepted but has no effect (it cannot). Expression arguments work — SUM(amount_cents + shipping_cents), COUNT(DISTINCT LOWER(customer)). Always give an aggregate an AS alias. Without one the output column is auto-named from the expression — sum(amount_cents), count(*) — which is awkward to index in every client language. HAVING compares an aggregate or a grouped column against a literal, combined with AND / OR. Comparing one aggregate against another is refused, and HAVING without GROUP BY is refused. SUM and AVG over a date, time or timestamp column are refused — the same way PostgreSQL refuses them. aggregates are the one shape that streams An aggregate directly over a scan or a filtered scan is computed as rows arrive, so SELECT COUNT(*) or SUM(...) over a very large table works without buffering it. That is not true of a projection — see limits. 11 Joins. INNER, LEFT, RIGHT and FULL OUTER all work, plus a two-table CROSS JOIN. A comma join with a single equality in the WHERE is promoted to an inner join for you. cURL Python TypeScript Go INNER JOIN # One equality per ON clause, both sides qualified: alias.column. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT o.id, c.name, o.amount_cents FROM shop.orders o INNER JOIN shop.customers c ON o.customer = c.id WHERE o.status = '\''paid'\'' LIMIT 10" }' # Projected columns keep their alias prefix: the result columns are # literally "o.id", "c.name", "o.amount_cents". # # JOIN is included from the Thunder configuration up - see section 14.result = db.sql.query(""" SELECT o.id, c.name, o.amount_cents FROM shop.orders o INNER JOIN shop.customers c ON o.customer = c.id WHERE o.status = 'paid' LIMIT 10 """) for row in result.rows: print(row["o.id"], row["c.name"]) # note the alias-qualified keysconst res = await db.sql(` SELECT o.id, c.name, o.amount_cents FROM shop.orders o INNER JOIN shop.customers c ON o.customer = c.id WHERE o.status = 'paid' LIMIT 10 `); if (res.kind === "select") { for (const row of res.rows as Record[]) { console.log(row["o.id"], row["c.name"]); // alias-qualified keys } }res, err := db.SQL(ctx, ` SELECT o.id, c.name, o.amount_cents FROM shop.orders o INNER JOIN shop.customers c ON o.customer = c.id WHERE o.status = 'paid' LIMIT 10 `) if err != nil { /* handle */ } for _, row := range res.Rows { fmt.Println(row["o.id"], row["c.name"]) // alias-qualified keys } one equality per ON clause — no AND, no inequality An ON clause must be exactly alias.column = alias.column. A composite join condition — ON a.x = b.x AND a.y = b.y — is refused, and so is any non-equality such as ON a.t > b.t. If you need a composite key, join on one column and filter the rest in the WHERE. The same limit means JOIN … USING (a, b) with more than one column is refused, though a single-column USING works. Projected columns keep their alias. SELECT o.id gives you a column literally named o.id. Rename in your client if you need something else. Up to 32 tables in one FROM; join chains are evaluated left to right. A JOIN with no ON is refused, and an alias cannot be reused across joins. GROUP BY, DISTINCT, ORDER BY and OFFSET all compose over a join. Window functions do not — that combination is refused. a typo on the right-hand table returns zero rows, not an error Column names are validated against the left-most table's manifest only. A misspelled column on a joined-in table passes validation and never matches, so the query succeeds with an empty result. If a join returns nothing and you expected rows, check the spelling on the right-hand side first. avoid NATURAL JOIN It is accepted, but it infers the join keys from the left table alone and trusts that the right table has them. When that assumption is wrong you get zero rows rather than an error. Write the ON clause out. 12 Writes and schema changes. Write statements execute — they are not translated into something you then have to re-issue. A successful INSERT, UPDATE or DELETE is durable when the response returns. cURL Python TypeScript Go INSERT with RETURNING # INSERT executes and writes durably. The column list is MANDATORY. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "INSERT INTO shop.orders (id, customer, amount_cents, status, notes, placed_ms) VALUES ($1, $2, $3, $4, $5, $6) RETURNING id, status", "params": ["ord-3001", "cus-90", 2500, "pending", "", 1714478400000] }' # Explicit upsert instead: # INSERT INTO shop.orders (id, status) VALUES ('ord-3001', 'shipped') # ON CONFLICT (id) DO UPDATE SET status = EXCLUDED.statusres = db.sql.execute(""" INSERT INTO shop.orders (id, customer, amount_cents, status, notes, placed_ms) VALUES ('ord-3001', 'cus-90', 2500, 'pending', '', 1714478400000) """) print(res.kind) # -> insert # Bulk loading? Use the row batch endpoint instead - it is far faster. db.rows.put_batch("shop.orders", rows)const res = await db.sql(` INSERT INTO shop.orders (id, customer, amount_cents, status, notes, placed_ms) VALUES ('ord-3001', 'cus-90', 2500, 'pending', '', 1714478400000) `); if (res.kind === "insert") { console.log(res.schema, res.rows); }res, err := db.SQL(ctx, ` INSERT INTO shop.orders (id, customer, amount_cents, status, notes, placed_ms) VALUES ('ord-3001', 'cus-90', 2500, 'pending', '', 1714478400000) `) if err != nil { /* handle */ } fmt.Println(res.Kind, res.Schema) // -> insert shop.ordersresponse200{ "kind": "insert", "schema": "shop.orders", "inserted": 1, "returning": ["id", "status"], "rows": [ { "id": "ord-3001", "status": "pending" } ] }response409{ "error": "constraint_violation", "detail": "..." }A duplicate primary key is a hard error - the existing row survives: cURL Python TypeScript Go UPDATE and DELETE # UPDATE executes. A WHERE clause is MANDATORY, and you cannot SET a # primary-key column. RETURNING works on UPDATE. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{"sql":"UPDATE shop.orders SET status = $1 WHERE customer = $2", "params":["shipped","cus-12"]}' # DELETE executes too, and DOES support RETURNING. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{"sql":"DELETE FROM shop.orders WHERE status = $1 RETURNING id", "params":["cancelled"]}'res = db.sql.execute( "UPDATE shop.orders SET status = 'shipped' WHERE customer = 'cus-12'" ) print(res.kind) # -> update # NOTE: res.rows_affected on this client is a placeholder, not the real # count. Read the raw response if you need it. db.sql.execute("DELETE FROM shop.orders WHERE status = 'cancelled'")await db.sql("UPDATE shop.orders SET status = 'shipped' WHERE customer = 'cus-12'"); const del = await db.sql("DELETE FROM shop.orders WHERE status = 'cancelled'"); if (del.kind === "delete") console.log(del.schema, del.pk);if _, err := db.SQL(ctx, "UPDATE shop.orders SET status = 'shipped' WHERE customer = 'cus-12'", ); err != nil { /* handle */ } del, err := db.SQL(ctx, "DELETE FROM shop.orders WHERE status = 'cancelled'") if err != nil { /* handle */ } fmt.Println(del.Kind, del.Schema)response{ "kind": "update", "schema": "shop.orders", "rows_affected": 2 }response{ "kind": "delete", "schema": "shop.orders", "returning": ["id"], "rows": [ {"id": "ord-1900"} ], "rows_affected": 1 } Every write statement, and what it does Statement Status Notes INSERT … VALUES (…) executes Multi-row VALUES supported. The column list is mandatory. A duplicate primary key is a 409, not an overwrite. INSERT … SELECT … executes Full source-scan authorization applies. INSERT … ON CONFLICT executes DO NOTHING and DO UPDATE SET col = literal | EXCLUDED.col. A conflict target is required; primary-key columns cannot be SET. INSERT … RETURNING executes Returns the written rows projected to the listed columns. UPDATE … SET … WHERE … executes WHERE is mandatory. Primary-key columns cannot be SET. UPDATE … FROM and joined UPDATE are refused. UPDATE … RETURNING executes Returns the updated rows with their NEW (post-SET) values. DELETE FROM … WHERE … executes Both the primary-key fast path and an arbitrary predicate. DELETE … RETURNING executes Returns the deleted rows. DELETE FROM t (no WHERE) executes Deletes every row. Accepted only on this endpoint — every other surface refuses a bare DELETE. CREATE TABLE · CREATE INDEX · CREATE VIEW · CREATE SEQUENCE · CREATE SCHEMA executes CREATE INDEX backfills existing rows before it returns. ALTER TABLE ADD / DROP / RENAME COLUMN · ADD / DROP CONSTRAINT executes Driven to completion synchronously — the change is live when the response returns. DROP TABLE · DROP VIEW · DROP SEQUENCE · DROP SCHEMA executes DROP TABLE is a full destructive purge: rows, indexes, relations and the registration. DROP SCHEMA is RESTRICT by default — it refuses while the namespace still owns tables; CASCADE drops every table in it. CREATE / DROP PROCEDURE · FUNCTION · CALL executes Preview scope — a single statement body for procedures, a scalar expression for functions. BEGIN · COMMIT · ROLLBACK executes Buffers writes into a session transaction. See the transactions page. Anything else refused Returns 400 listing the accepted statement verbs. a bare DELETE deletes everything, and this endpoint allows it DELETE FROM shop.orders with no WHERE is accepted here and removes every row. Every other surface refuses it. UPDATE is the opposite — it requires a WHERE as a safety check. Do not rely on the asymmetry; put a predicate on both. SQL INSERT rejects duplicates — the row endpoint does not An INSERT whose primary key already exists returns 409 and writes nothing; the existing row survives. That check covers both committed rows and duplicates inside the same VALUES list. The row endpoint is an upsert and overwrites instead — a real difference between the two write paths, worth knowing before you pick one. use the row batch endpoint for bulk loading Multi-row INSERT … VALUES works, but the batch row endpoint is the fast path by a wide margin, and it has a streaming form with no size limit. Reserve SQL INSERT for writes where you want the strict duplicate check or a RETURNING clause. Stored procedures and functions Preview scope. A procedure wraps one write statement whose parameters bind as $1..$N; run it with CALL. A function returns one value from a scalar expression and is usable inside a query. CREATE OR REPLACE is not supported - drop and recreate to change a definition. -- A procedure: one statement, parameters bound as $1..$N. CREATE PROCEDURE add_customer(id TEXT, email TEXT) AS BEGIN INSERT INTO shop.customers (id, email) VALUES ($1, $2) END; -- Invoke it with literal arguments. CALL add_customer('c_501', 'ada@example.com'); -- A scalar function returns one value from an expression. CREATE FUNCTION shout(s TEXT) RETURNS TEXT RETURN UPPER(s); -- Use it inside a query like any built-in. SELECT id, shout(email) FROM shop.customers; -- Change one by dropping and recreating (no CREATE OR REPLACE). DROP PROCEDURE add_customer; DROP FUNCTION shout; The procedure body is a single SELECT / INSERT / UPDATE / DELETE / CALL. Multi-statement bodies, procedural control flow, and the USING / DETERMINISTIC / REMOTE clauses are refused in preview. A function can also return a set with RETURNS TABLE(...). 13 EXPLAIN. Prefix any SELECT with EXPLAIN to get the plan back instead of the rows. The one thing to look for is the leaf: an index scan or index range scan means your predicate is using an index; a plain scan under a filter means it is reading the whole table. cURL EXPLAIN and EXPLAIN ANALYZE # Check whether an index is actually being used. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{"sql":"EXPLAIN SELECT id FROM shop.orders WHERE status = '\''paid'\''"}' # EXPLAIN ANALYZE runs the query and adds per-operator timings. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" -H "Content-Type: application/json" \ -d '{"sql":"EXPLAIN ANALYZE SELECT id FROM shop.orders WHERE status = '\''paid'\''"}'response{ "kind": "explain", "plan": "IndexScan shop.orders by_status ..." }response{ "kind": "explain", "plan": "...", "stats": { ... } } Predicate pushdown follows a fixed order: an equality on a single-column indexed column becomes an index scan; a >, < or BETWEEN on one becomes an index range scan; anything else becomes a full scan with a row-by-row filter. Predicate terms the index cannot serve are re-applied above it. 14 Where GROUP BY, JOIN and HAVING run. Projections, filters, sorting, limits, subqueries, set operations and plain aggregates all run on every configuration. GROUP BY, any JOIN and HAVING are bundled into every paid configuration from Thunder up, so there is no separate add-on to buy. On a configuration that does not include it, the request comes back 402 with a structured body naming what to enable: { "error": "addon_required", "addon": "sql-pro", "name": "SQL Pro", "purchase_url": "https://app.originchain.ai/billing/addons?enable=sql-pro", "msg": "This endpoint requires the SQL Pro add-on. Enable it at /app/billing/addons or have an admin do so." } Note that SELECT COUNT(*) FROM shop.orders is not gated — it is an aggregate without a GROUP BY. If you hit this on a configuration that should include it, check Billing → Add-ons in the console. the gate reads the statement text, not the parse tree It is a whole-word, case-insensitive scan for GROUP, JOIN, HAVING, LEFT, RIGHT, FULL and OUTER, and it runs before the statement is parsed — so it does not know a string literal from a keyword. WHERE name = 'GROUP' triggers it, and so does a column or table called left, outer or join. If you get an unexpected 402 on a query with no join in it, that is why: bind the literal as a parameter, or rename the column. 15 Limits & gotchas. the big one — a projection over a large table can fail with 413 Results are capped per query, by row count and by size. Past the cap the request fails outright: 413 { "error": "result_too_large", "unit": "rows", "observed": 240000, "cap": 200000, "msg": "The query's result set exceeds the engine's per-query memory cap. Add a LIMIT, narrow the filter, or page the ..." } The cap scales with the instance's memory — a few hundred thousand rows or a few hundred megabytes, whichever binds first. It is enforced during execution, so an unfiltered scan aborts partway rather than running to completion and then failing. Always put a LIMIT on an exploratory projection, and page large exports with OFFSET. ORDER BY … LIMIT is not a top-N stream Filters, projections and limits over a plain scan stream row by row, and so do aggregates. A sort does not — it materialises its whole input first, and so do joins, set operations and window functions. That means SELECT … ORDER BY x LIMIT 10 over a very large table buffers the entire table before it takes ten rows, and can hit the cap above even though you asked for ten rows. Narrow it with a WHERE on an indexed column first. Limit Value Notes Request body 8 MiB A statement larger than this is a 413. Bind parameters rather than inlining a large literal list. Statements per request 1 Semicolon-separated batches are refused. Tables per FROM 32 Over the cap the statement is refused with a message naming the limit. Subquery nesting 1 level A subquery inside a subquery is refused. Recursive CTE depth 100 The iteration cap on WITH RECURSIVE. Concurrent heavy queries scales with memory Over the limit a query waits briefly, then sheds with 429 and a Retry-After. Cross-shard result gather 1,000,000 rows On a sharded instance only. Over it, the query is refused with 501 — narrow it. Smaller edges worth knowing SELECT 1 fails. Every statement that names a column must name a table. Drivers and connection pools that probe liveness with a bare SELECT 1 will get an error — point them at a real table, or use the health endpoint. A constant-only expression such as SELECT 1 + 2 AS three does work. No NULLS FIRST / NULLS LAST. Missing values sort as JSON null. If null placement matters, add a CASE expression to the projection and sort on that. ORDER BY takes bare column names. Not ORDER BY LOWER(name), not ORDER BY a + b. Project the expression with an alias and order by that alias. A projection position works too, except over SELECT *. No derived tables. FROM (SELECT …) t is refused — use a WITH clause, which is supported. NOT BETWEEN is refused. Write col < a OR col > b — the error message on this one contains stale advice claiming OR is unavailable; it is available. A CTE's columns are not validated. Referring to a column the CTE does not actually produce is accepted and yields nulls rather than an error. OFFSET without LIMIT disables limit pushdown into an index range scan. Pair them when you can. Inside a session transaction, statements do not see the buffer. A SELECT after a buffered INSERT will not find the new row. See transactions. ← Docs home Next: vector search →`` --- # Constraints on OriginChainDB - FK + CHECK Canonical source: https://originchaindb.com/docs/sql/constraints Sitemap last modified: 2026-09-13T20:18:50.000Z docs · sql · constraints # FK + CHECK constraints. Foreign keys and CHECK constraints are declared in the TOML manifest and enforced at write time. Foreign keys ship with three on-delete actions — `no_action`, `restrict` and `set_null` — and CHECK follows SQL’s 3-valued logic. A violation comes back as HTTP 409 with a typed `constraint_violation` body: no silent skip, no partial commit. ## Foreign keys. ``` # manifest.toml namespace = "shop" table = "orders" primary_key = ["id"] # ... [[columns]] declarations ... [[foreign_keys]] name = "fk_client" from_col = "client_id" target_schema = "shop.clients" target_col = "id" on_delete = "restrict" # no_action | restrict | set_null ``` actionstatusbehavior `no_action` `shipped` Reject the delete if any child row references the parent. Same semantic as 'restrict' here. `restrict` `shipped` Reject the delete with a typed FK error if any child row references the parent. `set_null` `shipped` Set the referencing column to NULL in every child row, then delete the parent. `cascade` `deferred` Recursively delete child rows. Roadmap. `set_default` `deferred` Set the referencing column to its declared default. Roadmap. ## CHECK expressions. CHECK expressions support literals, `= != < <= > >=`, `AND / OR / NOT`, `IS [NOT] NULL`, and `IN (..)`. ``` [[check_constraints]] name = "amount_positive" expression = "amount >= 0 AND status IN ('pending', 'paid', 'shipped')" [[check_constraints]] name = "refund_window" expression = "refunded_at IS NULL OR refunded_at >= shipped_at" ``` ## 3-valued logic. CHECK uses the same semantic SQL has shipped for decades: any operand that is NULL makes the whole expression NULL, and a NULL expression passes the CHECK (it does not fail). That matches the industry-standard 3-valued-logic behaviour. If you want NULL to fail the CHECK, add an explicit `IS NOT NULL` guard. ## Errors. Constraint violations surface as typed HTTP errors: `409 Conflict` with a `constraint_violation` body containing the constraint name, kind (`foreign_key` or `check`), and the rejected row's primary key. No silent skip, no partial commit. --- # SQL examples by industry - orders, ledger - OriginChainDB Canonical source: https://originchaindb.com/docs/sql/industries Sitemap last modified: 2026-09-13T20:18:50.000Z sql · by industry # SQL models by industry Four models teams actually build: an order book, a financial ledger, inventory with reservations and an event stream. Each starts with the tables, because the shape of the tables decides which questions stay cheap. ## Retail: an order book Orders, their lines, and the products they point at. The reporting questions are all aggregates over a join. ### The tables ``` shop.products (id PK, name, brand, price) shop.orders (id PK, customer_id, status, placed_at) shop.lines (id PK, order_id -> orders.id, product_id -> products.id, qty) ``` ### Revenue by brand, last month One statement, joined and grouped in the engine, rather than three round trips and a loop in your service. ``` SELECT p.brand, SUM(l.qty * p.price) AS revenue, COUNT(DISTINCT o.id) AS orders FROM shop.lines l JOIN shop.orders o ON l.order_id = o.id JOIN shop.products p ON l.product_id = p.id WHERE o.status = 'paid' GROUP BY p.brand ORDER BY revenue DESC ``` ### The same catalog, searched The rows you just queried are the rows full-text and vector search see. Nothing is copied, so a product added by an INSERT is searchable immediately — see [the search quickstart](https://originchaindb.com/docs/elasticsearch/quickstart). ``` SELECT id, name FROM shop.products WHERE brand = 'Aero' ORDER BY price ASC ``` ## Financial services: a ledger Append-only entries, balances derived rather than stored, and money kept in minor units so totals never drift. ### The tables ``` fin.accounts (id PK, holder, currency, opened_at) fin.entries (id PK, account_id -> accounts.id, amount_minor, kind, booked_at) ``` ### Balance per account Derived from the entries, so it cannot disagree with them. ``` SELECT a.id, a.holder, a.currency, SUM(e.amount_minor) AS balance_minor FROM fin.accounts a JOIN fin.entries e ON e.account_id = a.id GROUP BY a.id, a.holder, a.currency ``` who sees which rows Row-level security and column masking are enforced by the engine, not by your query. A caller scoped to one desk gets that desk's rows from this exact statement, and a masked column stays masked in an aggregate. You do not add a filter for it, and you cannot forget to. ## Logistics: inventory with reservations Stock on hand, minus what is promised. The trap is reading availability and acting on it after someone else already did. ### The tables ``` wh.stock (sku PK, on_hand) wh.reservations (id PK, sku -> stock.sku, qty, state, made_at) ``` ### What is actually available Subtract live reservations from stock in one statement. ``` SELECT s.sku, s.on_hand, COALESCE(SUM(r.qty), 0) AS reserved, s.on_hand - COALESCE(SUM(r.qty), 0) AS available FROM wh.stock s LEFT JOIN wh.reservations r ON r.sku = s.sku AND r.state = 'held' GROUP BY s.sku, s.on_hand ``` make the decision atomic Reading availability and then reserving is two steps, and another writer fits between them. Do the read and the write in one transaction — see [transactions](https://originchaindb.com/docs/transactions) — so the reservation either wins or fails cleanly. ## Product analytics: an event stream High write volume, and questions that are almost always time-bucketed. ### The tables ``` app.events (id PK, user_id, name, props_json, at) ``` ### Daily active users Grouped in the engine over the whole table, not sampled in a job. ``` SELECT DATE_TRUNC('day', at) AS day, COUNT(DISTINCT user_id) AS dau FROM app.events WHERE at >= '2026-09-01T00:00:00Z' GROUP BY day ORDER BY day ``` when a query becomes a dashboard A query a dashboard runs every minute is a materialized view. Declare it once and read it as a table — see [materialized views](https://originchaindb.com/docs/sql/materialized-views). ## What these have in common - The tables are the design. Every cheap question later comes from a key or a foreign key you declared now. - One store. These same rows answer full-text, vector, graph and natural-language queries. Nothing is exported to a warehouse to be asked a question. - Policy is enforced below your query. Row-level security and masking apply to the statement itself, including inside aggregates. [Start from zero Schema, insert, select, join, aggregate.](https://originchaindb.com/docs/sql/quickstart)[More examples Single-purpose recipes with their responses.](https://originchaindb.com/docs/examples/sql) --- # Materialized views on OriginChainDB - install, refresh Canonical source: https://originchaindb.com/docs/sql/materialized-views Sitemap last modified: 2026-09-13T20:18:50.000Z docs · sql · materialized views # Materialized views. A materialized view is a precomputed query result that you can read like a table. OriginChainDB ships on-demand refresh: you install a view with a SQL definition, refresh it when you want the numbers updated, and read it like any other table. Incremental refresh stays on the roadmap. ## Three endpoints. - `POST :name/install` Define the view + compute the initial materialization in one call. - `POST :name/refresh` Re-run the definition and atomically overwrite the materialization. - `GET :name` Read the materialization. Constant-time vs the underlying GROUP BY. ## Install. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/sql/materialized_views/customer_revenue/install" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sql": "SELECT customer_id, SUM(qty * price) AS revenue FROM line_items GROUP BY customer_id" }' ``` ## Refresh. The refresh atomically overwrites the existing materialization. Readers see either the previous version or the new one — never a half-written state. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/sql/materialized_views/customer_revenue/refresh" \ -H "Authorization: Bearer $OC_TOKEN" ``` ## Read. #### cURL ``` curl "https://acme.db.originchain.ai/v1/tenants/$T/sql/materialized_views/customer_revenue" \ -H "Authorization: Bearer $OC_TOKEN" ``` ## Cost considerations. Refresh re-runs the full definition. Cost scales with the underlying query; pick refresh cadence to match how stale your readers can tolerate. For hot rollups (BI dashboards, leaderboards) refresh on a fixed cron — for cold rollups, refresh on demand. Incremental refresh (delta-only) is on the roadmap; until then, a full refresh is the only path. --- # SQL quickstart - your first query on OriginChainDB Canonical source: https://originchaindb.com/docs/sql/quickstart Sitemap last modified: 2026-09-23T08:33:23.000Z sql · quickstart # Run your first SQL query Create a table, insert sample rows, and run your first SQL query against OriginChainDB. ## Before you start An instance and an API key. Everything below is one endpoint: ``` export OC_URL='https://' export OC_TENANT='' export OC_TOKEN='' ``` ## The five steps 1. 1. Register a schema A schema declares a table and its columns. Register it once; every surface reads the same declaration. Full reference in [the schema reference](https://originchaindb.com/docs/schemas/reference). ``` curl -X POST "$OC_URL/v1/tenants/$OC_TENANT/schemas" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: text/plain" \ --data-binary @- <<'TOML' namespace = "shop" table = "products" primary_key = ["id"] [[columns]] name = "id" ty = "str" required = true [[columns]] name = "name" ty = "str" [[columns]] name = "price" ty = "i64" TOML ``` 2. 2. Insert a row The row endpoint is the fast path for writes and the one to use for ingest. INSERT over SQL works too — see step five. ``` curl -X POST "$OC_URL/v1/tenants/$OC_TENANT/rows/shop.products" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"id":"sku-8842","name":"Carbon Marathon","price":149}' ``` 3. 3. Select it back One endpoint, one JSON field. The response carries a kind so your code can switch on the statement type. ``` curl -X POST "$OC_URL/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sql":"SELECT id, name, price FROM shop.products WHERE price < 200"}' ``` 4. 4. Group and aggregate Aggregates, GROUP BY and HAVING run in the engine, not in your application. ``` SELECT brand, COUNT(*) AS n, AVG(price) AS avg_price FROM shop.products GROUP BY brand HAVING COUNT(*) > 2 ORDER BY avg_price DESC ``` 5. 5. Join two tables Joins resolve across schemas in the same tenant. Writes over SQL (INSERT, UPDATE) execute against the engine as well; for bulk ingest the row endpoints are still faster. ``` SELECT o.id, p.name, o.qty FROM shop.orders o JOIN shop.products p ON o.product_id = p.id WHERE o.status = 'paid' ``` ## From a PostgreSQL client The same data is reachable over the PostgreSQL wire protocol, so psql, DBeaver, pgAdmin and standard drivers connect directly. See [connect a SQL client](https://originchaindb.com/docs/connect-sql-client) for the connection string and the compatibility matrix. ## What to read next [SQL reference Every clause that works today, and the ones that do not, with the response shape per statement kind.](https://originchaindb.com/docs/sql)[Constraints Primary keys, foreign keys and what the engine enforces on write.](https://originchaindb.com/docs/sql/constraints)[Materialized views Precompute a query and keep it current.](https://originchaindb.com/docs/sql/materialized-views)[By industry Worked models: orders, ledgers, inventory, events.](https://originchaindb.com/docs/sql/industries) --- # Transactions - OriginChainDB atomic multi-statement writes Canonical source: https://originchaindb.com/docs/transactions Sitemap last modified: 2026-09-13T20:18:50.000Z reference · transactions # Transactions A transaction groups multiple SQL statements into one atomic unit: either every write lands, or none do. Transactions run over the same `POST /v1/tenants/:t/sql` endpoint you already use - and they may span tables anywhere in your configuration, including tables that live on different nodes. Commit is atomic either way. ## The lifecycle. Send `BEGIN`, then your statements, then `COMMIT` (or `ROLLBACK`) - each as its own `POST /sql` call. What ties them together is the `X-OC-Session-Id` header: every statement carrying the same session id belongs to the same transaction. Any unique string works; generating a UUID per transaction is the simple, collision-free choice. #### cURL ``` # Every statement goes to the same endpoint. A transaction is a # sequence of statements that share one X-OC-Session-Id header. SID=$(uuidgen) # any unique string; a UUID per transaction is the easy choice # 1. Open the transaction curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{ "sql": "BEGIN" }' # 2. Run statements - they are part of the transaction because they # carry the same session id. Tables may live on different nodes. curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{ "sql": "INSERT INTO shop.orders (id, customer, amount_cents) VALUES ('\''o-1'\'', '\''c-9'\'', 4200)" }' curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{ "sql": "UPDATE shop.inventory SET reserved = 1 WHERE id = '\''sku-7'\''" }' # 3. Commit - everything above becomes visible at once, or nothing does curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/sql" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "X-OC-Session-Id: $SID" \ -H "Content-Type: application/json" \ -d '{ "sql": "COMMIT" }' ``` Nothing written inside the transaction is visible to other readers until `COMMIT` returns 200. If `COMMIT` succeeds, every statement's writes are durable and visible together. ## What's guaranteed. | Guarantee | What it means for you | | --- | --- | | atomic commit | All writes in the transaction become visible together, or none do. There is no state in which a reader sees half of them. | | spans nodes | On a [multi-node configuration](https://originchaindb.com/docs/multi-node), a transaction can touch tables that live on different nodes. The atomicity guarantee is identical - your code doesn't change. | | no partial writes | If anything goes wrong - a statement error, a lost transaction, a failed commit - zero writes from the transaction are applied. | | safe rollback | `ROLLBACK` is idempotent and always returns 200 - even if the transaction no longer exists. It is always safe to call in a cleanup path. | ## Errors inside a transaction. If a statement inside a transaction fails - a constraint violation, an unknown column, a type mismatch - the transaction is poisoned. You can't paper over the error and keep going: a later `COMMIT` fails cleanly, and no writes from the transaction are applied. The correct response is to `ROLLBACK`, fix the statement, and run the whole transaction again. Separately, a transaction can be lost - for example when the service restarts while your transaction is open. In that case the next statement you send returns `409` with `"transaction not found"`. Nothing was committed. Treat 409 as a signal to `ROLLBACK` (safe, idempotent) and retry the entire transaction from `BEGIN` with a fresh session id. | You see | What happened | Do this | | --- | --- | --- | | 4xx on a statement | The statement failed; the transaction is poisoned. COMMIT will now fail cleanly. | ROLLBACK, fix the statement, retry the whole transaction. | | 409 "transaction not found" | The open transaction was lost (e.g. a service restart, or the idle timeout below). Nothing was committed. | ROLLBACK, then retry the whole transaction with a fresh session id. | | COMMIT fails | The transaction could not be applied. No partial writes exist. | ROLLBACK, retry the whole transaction. | ## Idle timeout. An open transaction that sits idle for 5 minutes is discarded. The next statement on that session id returns `409 "transaction not found"`, and nothing was committed. Transactions are for grouping writes, not for holding application state - keep them short, and don't hold one open across user think-time. ## The retry pattern. Because a lost or poisoned transaction never leaves partial writes behind, the safe recovery is always the same: `ROLLBACK`, then re-run the whole transaction. Wrap it once and reuse it: #### Python ``` import uuid, requests def run_transaction(base, tenant, token, statements, max_retries=3): """Run a list of SQL statements as one atomic transaction. Retries the WHOLE transaction from BEGIN if it is lost mid-flight.""" for attempt in range(max_retries + 1): sid = str(uuid.uuid4()) # fresh session id per attempt headers = { "Authorization": f"Bearer {token}", "X-OC-Session-Id": sid, } url = f"{base}/v1/tenants/{tenant}/sql" def sql(stmt): return requests.post(url, headers=headers, json={"sql": stmt}) sql("BEGIN") ok = True for stmt in statements: r = sql(stmt) if r.status_code == 409: # transaction not found - it was lost ok = False break r.raise_for_status() # statement errors poison the tx if ok: commit = sql("COMMIT") if commit.status_code == 200: return commit.json() sql("ROLLBACK") # always safe - idempotent, returns 200 if attempt == max_retries: raise RuntimeError("transaction failed after retries") ``` #### TypeScript ``` async function runTransaction( base: string, tenant: string, token: string, statements: string[], maxRetries = 3, ) { for (let attempt = 0; attempt <= maxRetries; attempt++) { const sid = crypto.randomUUID(); // fresh session id per attempt const sql = (stmt: string) => fetch(`${base}/v1/tenants/${tenant}/sql`, { method: "POST", headers: { "Authorization": `Bearer ${token}`, "X-OC-Session-Id": sid, "Content-Type": "application/json", }, body: JSON.stringify({ sql: stmt }), }); await sql("BEGIN"); let ok = true; for (const stmt of statements) { const res = await sql(stmt); if (!res.ok) { ok = false; break; } // 409 or a statement error } if (ok) { const commit = await sql("COMMIT"); if (commit.ok) return commit.json(); } await sql("ROLLBACK"); // always safe - idempotent } throw new Error("transaction failed after retries"); } ``` common mistakes - Retrying only the failed statement. After a 409 the transaction is gone - the earlier statements in it were never committed. Retry from `BEGIN`, not from where you left off. - Reusing a session id across transactions. Generate a fresh id per transaction (and per retry attempt). Reuse invites cross-talk between logically separate transactions. - Holding a transaction open while waiting on something else. An LLM call, a user prompt, a slow upstream API - all of these can outlive the 5-minute idle timeout. Gather your data first, then run the transaction. - Skipping ROLLBACK because "it probably doesn't exist". ROLLBACK is idempotent and returns 200 either way. Always call it in your error path - it costs nothing and never hurts. --- # Troubleshooting - OriginChainDB common problems and fixes Canonical source: https://originchaindb.com/docs/troubleshooting Sitemap last modified: 2026-09-13T20:18:50.000Z reference · troubleshooting # Troubleshooting This page lists 23 problems as a symptom and its fix side by side, covering authentication, connection, schema registration, row inserts, SQL, vector, full-text, graph, rate limits and PITR restore. Search it (⌘F / Ctrl-F) for your exact error message. For error codes specifically, see [Error reference](https://originchaindb.com/docs/errors). For rate-limit issues, see [Rate limits](https://originchaindb.com/docs/rate-limits). If you don't find your problem here, email [support@originchain.ai](mailto:support@originchain.ai) with the `X-OC-Trace-Id` from the response headers. | Problem | Symptom + fix | | --- | --- | | 401 unauthorized on every call | Symptom: Bearer token rejected even though you copied it from the dashboard. Fix: Token may have been truncated (often by Slack or terminals trimming spaces). Open the dashboard's token panel and copy again. Also check the endpoint hostname matches the token's instance - tokens are per-instance. | | Cannot connect to host | Symptom: DNS or TLS errors before any request lands. Fix: Verify OC_HOST is the full hostname (e.g. acme.db.originchain.ai), not just the tenant id. The dashboard's 'Connect' panel shows the right URL. | | Schema registration returns 400 'unknown column type' | Symptom: Trying ty = 'ulid' / 'string' / 'int' - spellings borrowed from another database. Fix: `decimal`, `timestamp` and `vector` are all valid - it is the aliases that are not. The type names are `i64`, `u64`, `f64`, `bool`, `str`, `text`, `bytes`, `timestamp`, `date`, `time`, `uuid`, `decimal`, `enum`, `json`, `list`, `inet`, `interval`, `point` and `vector`. So `ulid` is `str`, `string` is `str` and `int` is `i64`. See Schemas reference. | | Schema registration returns 400 'invalid type: string, expected struct' | Symptom: Relation target was written as a single string, e.g. target = 'shop.suppliers'. Fix: target is a sub-table: [relations.target] namespace = '...' table = '...' pk = '...'. See Schemas → relations. | | [[extractions.fts]] and [[extractions.vector]] blocks ignored | Symptom: Register succeeds but vector / FTS data never appears. Fix: Those blocks aren't real - they don't exist in the schema grammar. FTS indexes via POST /v1/tenants/:t/fts/:table/:field. Vectors via POST /v1/tenants/:t/vector/:table/put. | | Row insert returns 400 type_mismatch | Symptom: Specific field named in the error message. Fix: Check that the JSON value's type matches the column's declared ty. Numbers can't be sent as JSON strings. Booleans need to be true / false, not 'true' / 'false'. | | Row insert returns 400 fk_violation | Symptom: Foreign-key column points at a row that doesn't exist. Fix: Insert the target row first. If you really want orphan references (unusual), drop the foreign-key constraint from the schema. | | Bulk insert is much slower than expected | Symptom: Single-row inserts in a loop instead of using _batch. Fix: Use POST /rows/:schema/_batch with a JSON array (up to ~8 MiB) or NDJSON (no cap, streamed). See Insert → bulk. | | Duplicate rows after a network retry | Symptom: Single-row insert succeeded server-side but the client thought it failed and retried. Fix: Pass an Idempotency-Key header on every mutating call (or use the SDKs - they do this automatically). The engine dedupes retries within 24h. | | HAVING or a window function over a JOIN returns 400 | Symptom: A grouped or windowed query runs on one table, then fails once a JOIN is added. Fix: Plain `HAVING` and window functions ship. What is still refused is the join-shaped combination: `HAVING` across a JOIN, a `DISTINCT` aggregate over a JOIN, and a window function in the same `SELECT` as a JOIN. Aggregate one side first, then join. See [SQL reference](https://originchaindb.com/docs/sql). | | Vector topk returns 400 dim_mismatch | Symptom: Query vector length doesn't match the table's locked dim. Fix: Every vector in a table has the same length, set by the first put. If you switched models, you need a new table. | | Vector topk returns 400 metric_mismatch | Symptom: Used a different distance metric than the table's first put. Fix: Pick cosine / dot / l2 / manhattan once per table. Cosine is the right default for most text embeddings. | | Vector search returns fewer than k hits | Symptom: Asked for k=50 but got back 12. Fix: Either the table has fewer than 50 vectors, or your metadata filter is highly selective. Both are expected behaviour - it's not a bug. | | FTS search returns no hits for words that should match | Symptom: Text was indexed; query has the right words. Fix: Likely an analyzer mismatch. Stemming / lowercase / accent folding must be the same at index and query time. Re-check the analyzer configured on the (table, field). | | FTS hits link to rows that don't exist | Symptom: Search returns doc_ids, but GET /rows/:schema/:doc_id is 404. Fix: FTS indexes are independent of rows. Either the row was deleted (didn't trigger an FTS removal) or the indexed doc_id was never a real row PK to begin with. Re-issue the index call with current doc_ids. | | Graph neighbors returns empty | Symptom: Schema has relation declared, rows are written, but graph endpoints return no neighbors. Fix: Confirm the relation's from_col matches the column actually populated on the row. Rows written before the relation was declared don't have edges - re-write them to materialize. | | BFS takes a long time / times out | Symptom: Calling /bfs without max_depth or with a very large depth. Fix: Always cap max_depth at the smallest value that gives useful results. On dense graphs, depth=5 can return millions of nodes. | | 429 rate_limited on legitimate traffic | Symptom: Burst traffic during a scheduled job. Fix: Honor Retry-After. For sustained high throughput, request a limit increase (support@originchain.ai). For one-time imports, use _batch instead of single-row inserts. | | Replication lag visible in the dashboard | Symptom: Reads occasionally see stale data after writes. Fix: Reads are routed to the primary by default - they see writes immediately. If you're explicitly reading from a replica, accept eventual consistency. For strict read-after-write, read from the primary endpoint. | | PITR restore is taking longer than expected | Symptom: Restore in progress, ETA unknown. Fix: PITR is bounded by your recovery-point retention window (typically 7-30 days). Larger windows take longer to scan. Watch progress in the dashboard's Restore panel. | --- # Vector search on OriginChainDB - HNSW, IVF, IVF-PQ Canonical source: https://originchaindb.com/docs/vector Sitemap last modified: 2026-09-23T08:33:23.000Z reference · vector # Vector search Find the nearest stored embeddings to a query vector, with metadata filters and an index suited to your workload. For how to save a vector, see [Insert → vector](https://originchaindb.com/docs/insert#vector). For how to declare vector columns on a schema, see [Schemas → vector fields](https://originchaindb.com/docs/schemas/vector). All examples assume you have a client set up - see [Quickstart](https://originchaindb.com/docs/quickstart). ## 1. Top-k search. what this does Given a query embedding, return the top `k` rows whose stored embeddings are closest to it. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.011, -0.082, 0.046, /* ... 768 floats ... */], "k": 10, "dim": 768, "metric": "cosine" }' ``` #### Python ``` # query_768d is your query embedding - any list of 768 floats. hits = db.vector_topk( "shop.products", query=query_768d, k=10, dim=768, metric="cosine", ) for hit in hits: print(hit.id, hit.score) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, // number[] of length 768 k: 10, dim: 768, metric: "cosine", }); for (const hit of hits) { console.log(hit.id, hit.score); } ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, // []float32 of length 768 K: 10, Dim: 768, Metric: "cosine", }) if err != nil { /* handle */ } for _, h := range hits { fmt.Println(h.ID, h.Score) } ``` what each field means | Field | Type | Required | What it is | | --- | --- | --- | --- | | query | float[] | yes | The query embedding. Length must match the column's `dim`. | | k | int | yes | How many results to return. Typical values: 10, 50, 100. | | dim | int | yes | The vector's length. Must match the table's configured dimension. | | metric | string | no | How "closeness" is measured. Must match what was used at insert time. See [distance metric](https://originchaindb.com/docs/vector#metric). | | filter | object | no | Metadata filter. See [filter by metadata](https://originchaindb.com/docs/vector#filtered). | | mode | string | no | `"high_recall"` (default) or `"fast"`. See [speed vs recall](https://originchaindb.com/docs/vector#mode). | what you get back An array of `{ id, score }` objects, ordered from closest to farthest. The `id` is the row's primary key (so you can look up the full row); `score` tells you how close the match is. ``` [ { "id": "sku-9281", "score": 0.92 }, { "id": "sku-4017", "score": 0.88 }, { "id": "sku-3320", "score": 0.85 }, ... ] ``` Score ordering depends on the metric. With cosine and dot, higher is closer. With L2, lower is closer. Results always come back in correct ranking order - you do not need to sort yourself. common mistakes - Wrong dim. If your query vector is 1536 floats but the table was set up for 768, you get a `400` reading `vector has 1536 dims but collection "shop.orders" expects 768`. There is no `dim_mismatch` code - the body carries the message, not a token. - Wrong metric. Query with a different metric than the table was written under and you get a `409` `vector_metric_mismatch`, naming both metrics and a rebuild URL. If the collection declares a non-default `[vector].distance`, contradicting it is refused earlier, with a `400`. - Empty results from a brand-new table. Not an index lag. There is no build step for the default index - the embedding and the updated graph ship in one batch, so a vector is queryable as soon as the put returns. Fewer than k hits means fewer than k vectors match; retrying will not change it. ## 2. Filter by metadata. what this does Restrict the search to vectors whose metadata matches a filter. Useful for things like "find similar products but only in the shoes category". #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [/* 768 floats */], "k": 10, "dim": 768, "metric": "cosine", "filter": { "category": "running-shoes" } }' ``` #### Python ``` hits = db.vector_topk( "shop.products", query=query_768d, k=10, dim=768, metric="cosine", filter={"category": "running-shoes"}, ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, k: 10, dim: 768, metric: "cosine", filter: { category: "running-shoes" }, }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, K: 10, Dim: 768, Metric: "cosine", Filter: map[string]any{"category": "running-shoes"}, }) ``` Filters use exact equality on metadata fields you stored at insert time. The filter is applied during the search, not after - so a highly selective filter (e.g., only 1% of rows match) is still fast. common mistakes - Filtering on a field you didn't store. The filter looks at the `metadata` object you passed at insert time - not at the row's other columns. If you want to filter on a column, include it in metadata. - Range filters. Only exact equality is supported today (`category = "shoes"`). For range filters (`price < 100`), filter the result in your app after the search returns. - Very selective filters returning fewer than k hits. If only 5 rows match your filter, you get 5 hits even if you asked for 50. That's not a bug - it's all there is. ## 3. Distance metric. what this does Picks the math used to compare two vectors. You send the metric on each request - on the write and on the search - and the two have to use the same one. A collection can pin it once with `[vector].distance` on the schema, which then fills in an omitted `metric` and refuses one that contradicts it. pick by your model | Metric | Use it when | | --- | --- | | cosine | Default for text. Use this with OpenAI, Cohere, Voyage, BGE, E5, and most other text embedding models. Looks at the angle between two vectors - magnitude doesn't matter. | | dot | Use when your vectors are already unit-normalised (length 1). Slightly cheaper than cosine for the same result. Some image-embedding pipelines emit normalised vectors. | | l2 | Use when the absolute distance between vectors matters - some image-feature pipelines, certain audio embeddings, and a few specialty research models. | | manhattan | Same as L2 but uses absolute differences instead of squared ones. More forgiving of one or two big-difference dimensions. Niche. | Not sure? Pick `cosine`. It works for 90% of text embeddings. ## 4. Speed vs recall. what this does Approximate nearest-neighbor search trades a small amount of accuracy for huge speed gains. The `mode` field picks where on that trade-off you want to land. #### cURL ``` # Add "mode" to the topk body. "high_recall" (default) or "fast". curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [/* 768 floats */], "k": 10, "dim": 768, "metric": "cosine", "mode": "fast" }' ``` #### Python ``` hits = db.vector_topk( "shop.products", query=query_768d, k=10, dim=768, metric="cosine", mode="fast", # "fast" | "high_recall" (default) ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.products", { query: query768d, k: 10, dim: 768, metric: "cosine", mode: "fast", // "fast" | "high_recall" }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.products", originchain.VectorTopKRequest{ Query: query768d, K: 10, Dim: 768, Metric: "cosine", Mode: originchain.ModeFast, // ModeFast | ModeHighRecall }) ``` two modes | Mode | What you get | Use when | | --- | --- | --- | | high_recall (default) | ~96% of the truly-closest vectors. Higher latency. | Product search, similar-customer lookup, anywhere first-pass accuracy matters. | | fast | ~70% of the truly-closest vectors. ~3x faster. | RAG with a re-ranker, hot dashboards, anywhere latency dominates. | "Recall" means: of the truly closest vectors that brute force would return, how many did the index find? Both modes return ranking-correct results - the difference is whether the absolute top-K is occasionally missed. ## 5. Index choice. what this does OriginChainDB supports several different index types for vector data. Most users should stick with the default (HNSW). The other types are for very large or memory-constrained corpora. | Index | Pick when | | --- | --- | | HNSW (default) | Best accuracy, and the right choice for almost everyone. The engine puts it comfortably at around a million vectors per table at the default search width. | | IVF | Above ~10M vectors. Cheaper memory at the cost of a small recall hit. See [IVF reference](https://originchaindb.com/docs/vector/ivf). | | IVF-PQ | For the largest tables, or when memory is the constraint. Compresses vectors 64×-768×. See [IVF-PQ reference](https://originchaindb.com/docs/vector/ivf-pq). | | Binary quantization | When you need 32× memory savings and can tolerate the recall hit. See [Quantization reference](https://originchaindb.com/docs/vector/quantization). | | Sparse vectors | For models like SPLADE / uniCOIL that emit sparse vectors instead of dense ones. | The index is picked per request, not on the schema - `index` is a field on both the write and the query body, and the manifest has no `index` key. See [Index kinds](https://originchaindb.com/docs/vector#ex-index-kinds) for the accepted values. ## 6. Examples. Every operation below is shown in cURL, Python, TypeScript and Go. Where an SDK does not wrap an endpoint the tab says so and shows the raw call instead. 6.1 ## Insert vectors. `id`, `embedding` and `dim` are required. `metadata` is a free-form JSON object stored alongside the vector — it is what [filtering](https://originchaindb.com/docs/vector#topk-filter-shape) reads later, so put anything you may want to narrow on in there at write time. Writes are indexed eagerly and atomically: the embedding and the updated graph ship in one batch, so the index can never lag the data across a crash. There is no "build the index" step for the default index — a vector is queryable as soon as the call returns. single insert #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "embedding": [0.0124, -0.0883, 0.0451], "dim": 3, "metric": "cosine", "metadata": { "status": "paid", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001" } }' # → 201 Created, empty body ``` #### Python ``` db.vector_put( "shop.orders", id="01JTRX9KQ3YH8K2WMX0F5JZAB7", embedding=[0.0124, -0.0883, 0.0451], dim=3, metric="cosine", metadata={"status": "paid", "customer": "01JTRX1H4Q9P0N2WMX0F5JZ001"}, ) # The typed namespace infers dim from len(embedding) and sends no # metric - use it when cosine (the default) is what you want: db.vector.put( "shop.orders", "01JTRX9KQ3YH8K2WMX0F5JZAB7", [0.0124, -0.0883, 0.0451], metadata={"status": "paid"}, ) ``` #### TypeScript ``` await db.vectorPut("shop.orders", { id: "01JTRX9KQ3YH8K2WMX0F5JZAB7", embedding: [0.0124, -0.0883, 0.0451], dim: 3, metric: "cosine", metadata: { status: "paid", customer: "01JTRX1H4Q9P0N2WMX0F5JZ001" }, }); ``` #### Go ``` err := db.VectorPut(ctx, "shop.orders", originchain.VectorPutRequest{ ID: "01JTRX9KQ3YH8K2WMX0F5JZAB7", Embedding: []float32{0.0124, -0.0883, 0.0451}, Dim: 3, Metric: "cosine", Metadata: map[string]any{"status": "paid"}, }) ``` bulk insert Bulk is the right shape for ingest: it builds the graph in one pass and writes one log frame for the whole batch. `dim`, `metric`, `quantization` and `index` live on the envelope; each item carries only `id`, `embedding` and `metadata`. The body-size limit is lifted on this route only. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/put_bulk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "dim": 768, "metric": "cosine", "vectors": [ { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "embedding": [/* 768 floats */], "metadata": { "status": "paid" } }, { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB8", "embedding": [/* 768 floats */], "metadata": { "status": "refunded" } } ] }' ``` #### Python ``` # No SDK wraps bulk vector insert yet - call the endpoint directly. import httpx httpx.post( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.orders/put_bulk", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={ "dim": 768, "metric": "cosine", "vectors": [ {"id": "01JTRX...AB7", "embedding": vec_a, "metadata": {"status": "paid"}}, {"id": "01JTRX...AB8", "embedding": vec_b, "metadata": {"status": "refunded"}}, ], }, ) ``` #### TypeScript ``` // No SDK wrapper for bulk vector insert yet - using fetch directly. await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/vector/shop.orders/put_bulk`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ dim: 768, metric: "cosine", vectors: [ { id: "01JTRX...AB7", embedding: vecA, metadata: { status: "paid" } }, { id: "01JTRX...AB8", embedding: vecB, metadata: { status: "refunded" } }, ], }), }, ); ``` #### Go ``` // No SDK wrapper for bulk vector insert yet - using net/http directly. body, _ := json.Marshal(map[string]any{ "dim": 768, "metric": "cosine", "vectors": []map[string]any{ {"id": "01JTRX...AB7", "embedding": vecA, "metadata": map[string]any{"status": "paid"}}, {"id": "01JTRX...AB8", "embedding": vecB, "metadata": map[string]any{"status": "refunded"}}, }, }) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+"/vector/shop.orders/put_bulk", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") http.DefaultClient.Do(req) ``` response ``` { "inserted": 2, "elapsed_ms": 14 } ``` - At most 100,000 vectors per call — over that the engine answers `413`, not `400`. - An empty `vectors` array is a `200` with `inserted: 0` and no write. - Duplicate ids inside one batch are last-writer-wins. The whole batch is one atomic write. - Single insert returns `201 Created` with an empty body — don't try to parse it. sparse vectors Learned-sparse embeddings go to `POST /vector/:table/put_sparse` with `{ id, indices, values, dim, metadata? }`, and are queried with `POST /vector/:table/topk_sparse`. `indices` and `values` must be the same length, every index must be inside `dim`, and non-finite values are refused. No SDK wraps the sparse endpoints today. 6.2 ## Query — top-k. `query`, `k` and `dim` are required; everything else has a default. The field is `k` — there is no `top_k` alias. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.0124, -0.0883, 0.0451], "k": 5, "dim": 3, "metric": "cosine" }' ``` #### Python ``` hits = db.vector_topk( "shop.orders", query=[0.0124, -0.0883, 0.0451], k=5, dim=3, metric="cosine", ) for h in hits: print(h.id, h.score) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.orders", { query: [0.0124, -0.0883, 0.0451], k: 5, dim: 3, metric: "cosine", }); for (const h of hits) console.log(h.id, h.score); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.orders", originchain.VectorTopKRequest{ Query: []float32{0.0124, -0.0883, 0.0451}, K: 5, Dim: 3, Metric: "cosine", }) for _, h := range hits { fmt.Println(h.ID, h.Score) } ``` response ``` [ { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB7", "score": 0.9421 }, { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB8", "score": 0.9187 }, { "id": "01JTRX9KQ3YH8K2WMX0F5JZAB9", "score": 0.8804 } ] ``` A bare JSON array, not an envelope — no `hits` wrapper and no total count. Sorted by `score` descending, and larger always means closer: for `l2` and `manhattan` the distance is returned negated so the ordering convention holds across every metric. `k: 0` returns `[]` without touching storage. The ceiling on `k` is 4096 by default; above it you get a `400`. 6.3 ## Query modes — fast vs high_recall. `mode` is the one recall/latency knob exposed on the API. It sets the search beam width and nothing else — the build-time graph parameters are unaffected. | mode | Beam width | Measured (100k vectors, 128-dim) | | --- | --- | --- | | "high_recall" | 1200 | Default. recall@10 ≈ 0.96, p99 ≈ 109 ms | | "fast" | 300 | recall@10 ≈ 0.69, p99 ≈ 37 ms | #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [/* 768 floats */], "k": 10, "dim": 768, "metric": "cosine", "mode": "fast" }' ``` #### Python ``` # Only the legacy method carries `mode`; the typed # db.vector.topk() namespace does not accept it. hits = db.vector_topk( "shop.orders", query=query_768d, k=10, dim=768, metric="cosine", mode="fast", ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.orders", { query: query768d, k: 10, dim: 768, metric: "cosine", mode: "fast", }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.orders", originchain.VectorTopKRequest{ Query: query768d, K: 10, Dim: 768, Metric: "cosine", Mode: originchain.ModeFast, }) ``` - Omitting `mode` gives you `high_recall`. That is the safe default and the right one for first-pass retrieval. - The value is case-sensitive. An unknown value is a hard `400`: ``unknown `mode` "FAST": expected one of "fast", "high_recall"``. - `mode` applies to the default graph index only. On `ivf` and `ivf_pq` queries it is accepted and then ignored — those paths have no beam width. Use [`nprobe`](https://originchaindb.com/docs/vector#ivfpq-install) there instead. - The recall figures above are the published measurements for one corpus shape. Treat them as a guide to the trade-off, not an SLA for your data. seeing what a query actually did `POST /vector/:table/topk_explain` takes the same body and returns `{ hits, config }`, where `config` reports the resolved `metric`, `dim`, `k`, `ef_search`, `m` and `beam_width`. Useful for confirming a mode actually took effect. It always runs the graph index, whatever `index` you pass. 6.4 ## Metadata filtering. `filter` is a flat map of strict equalities against the `metadata` you stored at write time. Multiple keys are ANDed. There are no operators — no ranges, no `$in`, no nesting, no negation. #### cURL ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [/* 768 floats */], "k": 10, "dim": 768, "metric": "cosine", "filter": { "status": "paid" } }' ``` #### Python ``` hits = db.vector.topk( "shop.orders", query_768d, k=10, metric="cosine", filter={"status": "paid"}, ) ``` #### TypeScript ``` const hits = await db.vectorTopk("shop.orders", { query: query768d, k: 10, dim: 768, metric: "cosine", filter: { status: "paid" }, }); ``` #### Go ``` hits, err := db.VectorTopK(ctx, "shop.orders", originchain.VectorTopKRequest{ Query: query768d, K: 10, Dim: 768, Metric: "cosine", Filter: map[string]any{"status": "paid"}, }) ``` Values are compared with strict JSON equality, so `1` does not match `"1"` and `true` does not match `"true"`. A record that is missing a key named in the filter is rejected, not treated as null. a filtered query can return fewer than k hits On the default graph index, filtering is a post-filter: the engine searches for `k × 4` candidates and then drops the ones that don't match. If your filter is highly selective, most of the over-fetched set is discarded and you get back fewer than `k` results — even when more matching vectors exist. Ask for a larger `k` than you need when you filter narrowly. filter is silently ignored on ivf_pq A `filter` sent with `"index": "ivf_pq"` is not applied, and the request still returns `200` with unfiltered results. This is the sharpest edge on this page: filter and IVF-PQ do not compose today. Filter on the default index, or apply the predicate yourself after the hits come back. On `"index": "ivf"` the filter is honoured, and applied inside each cell scan before scoring. 6.5 ## Index kinds. `index` appears on both the write and the query body, and the two must agree — a vector written under one index kind is not visible to a query using another. Omitting it everywhere is the common and correct choice. | index | What it is | | --- | --- | | "hnsw" | The default. A navigable small-world graph, built incrementally on every insert. Build parameters are fixed: 16 neighbours per node (32 at the base layer) and a construction beam of 200. No setup, no training, no minimum corpus. This is what you want unless you have measured a reason otherwise. | | "ivf" | Inverted file: vectors are assigned to the nearest of K centroids, and a query scans only `nprobe` of those cells. Requires centroids to be installed or trained first. Filters are applied inside the cell scan. | | "ivf_pq" | Inverted file plus product quantization — each vector is stored as a short code instead of full floats. This is the memory-footprint option for large corpora, and the one the [presets](https://originchaindb.com/docs/vector#ivfpq-install) build. | ### Approximate, not exact. All three index kinds are approximate nearest-neighbour structures: they trade a small amount of recall for a large amount of speed, and none of them guarantees that the true nearest neighbour is in the result set. That is the deal ANN makes, and it is almost always the right one — an exhaustive scan is exact but linear in corpus size, which stops being viable long before a million vectors. If exactness genuinely matters for a small collection, raise `k` and re-rank the candidates yourself with your own distance function. 6.6 ## Building an IVF-PQ index from a preset. IVF-PQ has a lot of knobs — partition count, subspace count, code width. Rather than expose them, the engine ships exactly two named presets and derives every parameter from your corpus. `preset` is the only required field. #### cURL ``` curl -X POST \ "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/create-ivf-pq-index" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "preset": "compressed" }' ``` #### Python ``` # No SDK wraps the preset index build - call the endpoint directly. import httpx r = httpx.post( f"https://{OC_HOST}/v1/tenants/{OC_TENANT}/vector/shop.orders/create-ivf-pq-index", headers={"Authorization": f"Bearer {OC_TOKEN}"}, json={"preset": "compressed"}, timeout=None, # training reads the whole corpus ) print(r.json()["pq_m"], r.json()["partitions"]) ``` #### TypeScript ``` // No SDK wrapper for the preset index build - using fetch directly. const r = await fetch( `https://${process.env.OC_HOST}/v1/tenants/${process.env.OC_TENANT}/vector/shop.orders/create-ivf-pq-index`, { method: "POST", headers: { "Authorization": `Bearer ${process.env.OC_TOKEN}`, "Content-Type": "application/json", }, body: JSON.stringify({ preset: "compressed" }), }, ); const cfg = await r.json(); console.log(cfg.pq_m, cfg.partitions); ``` #### Go ``` // No SDK wrapper for the preset index build - using net/http directly. body, _ := json.Marshal(map[string]any{"preset": "compressed"}) req, _ := http.NewRequestWithContext(ctx, "POST", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/vector/shop.orders/create-ivf-pq-index", bytes.NewReader(body)) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) req.Header.Set("Content-Type", "application/json") resp, err := http.DefaultClient.Do(req) ``` response — the config it chose ``` { "trained": true, "installed": true, "preset": "compressed", "partitions": 1024, "pq_m": 48, "pq_bits": 8, "keep_raw": false, "dim": 768, "training_corpus_size": 50000 } ``` The optional `seed` field (default `0`) makes training deterministic. There are no other fields — you cannot override partitions, subspace count or code width. build then query — no re-insert needed This endpoint populates the index cells as part of the build, feeding your stored vectors back through the same write path an IVF-PQ insert would use. The index is queryable the moment the call returns — you do not have to re-insert your rows under `index: "ivf_pq"`. This is specific to this preset endpoint: the lower-level centroid-install primitives deliberately do not write postings. ### How each preset resolves. Both presets pick the same partition count and the same 8-bit codes. They differ in how finely the vector is subdivided, and in whether the original vector is kept. | Parameter | high_recall | compressed | | --- | --- | --- | | Target sub-vector width | 8 | 16 | | `pq_m` (subspaces) | The divisor of `dim` — capped at 64 — whose sub-vector width `dim / m` lands closest to the target above. | | | `pq_bits` | 8 | 8 | | `keep_raw` | true | false | | `partitions` | `4 × √N` rounded to the nearest power of two, clamped to `[64, 65536]`. Identical for both presets. | | Because `pq_m` must divide `dim` and is capped at 64, the two presets converge at high dimensionality. Worked from the real formula: | dim | high_recall pq_m | compressed pq_m | Code bytes / vector | | --- | --- | --- | --- | | 128 | 16 | 8 | 16 B vs 8 B | | 768 | 64 | 48 | 64 B vs 48 B | | 1536 | 64 | 64 | 64 B vs 64 B | At 1536 dimensions the codes are identical — both presets hit the 64-subspace cap — and the only remaining difference is `keep_raw`. Since `keep_raw` stores an extra `dim × 4` bytes per vector, at 1536 dimensions `high_recall` costs about 6.2 KB per vector against 64 bytes for `compressed` — roughly a 97× difference in footprint for the same search behaviour today. what keep_raw does and does not buy you today `high_recall` retains the full-precision vector so that a candidate set can later be re-scored exactly. The HTTP top-k path does not perform that re-rank today — an `ivf_pq` query scores against the quantized codes for both presets. So on the current API the two presets return comparable results, and `high_recall`'s extra storage is buying future re-ranking rather than present accuracy. If footprint is why you are reaching for IVF-PQ at all, `compressed` is the honest choice. ### Which to pick. - Neither, at small scale. Under a few hundred thousand vectors the default graph index is faster and more accurate, and needs no build step. IVF-PQ is a memory-footprint tool, not a speed tool. - `compressed` when the corpus no longer fits comfortably in memory and you want the smallest possible resident footprint. - `high_recall` when you want the raw vectors kept alongside the codes — for exact re-ranking you perform yourself, or to be ready for engine-side re-ranking without a rebuild. ### Querying the built index. Pass `"index": "ivf_pq"` and, optionally, `nprobe` — the number of cells to visit. ``` curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [/* 768 floats */], "k": 20, "dim": 768, "metric": "cosine", "index": "ivf_pq", "nprobe": 16 }' ``` - `nprobe` defaults to `ceil(sqrt(partitions))`. Higher visits more cells: better recall, more work. - Valid range is `1` to 256. `0` is a `400` (`nprobe must be >= 1`); above 256 is a `400` naming the cap. - `nprobe` is ignored on the default graph index — it only means something for `ivf` and `ivf_pq`. - In the Python SDK `nprobe` is available on the typed `db.vector.topk()` namespace. No other SDK exposes it. ### The minimum-corpus rule. Training refuses to run on a corpus too small to populate its partitions: you need at least four vectors per partition. Because the partition count is floored at 64, that means a practical floor of 256 vectors — and far more once `4 × √N` pushes the partition count up. ``` { "error": "not enough vectors to build a Compressed index: 1024 partitions need >=4*K = 4096 vectors, found 900" } ``` Building on an empty table is a separate `400`: `no vectors stored for this table; put vectors before building an index`. Note that the preset name is echoed in the error in its internal capitalised form (`Compressed`, `HighRecall`) rather than the wire form you sent. the lower-level centroid endpoints Plain `ivf` (without quantization) is driven by four separate endpoints, wrapped by the Python SDK only: - `POST …/train-and-install-centroids` — `db.vector.train_and_install_centroids(table, partitions=…)`. Same four-per-partition minimum. - `POST …/install-centroids` — `db.vector.install_centroids(table, centroids)` for centroids you trained elsewhere. - `GET …/centroids` — `db.vector.centroids(table)`, a truncated preview. - `GET …/ivf-rebalance-status` — `db.vector.rebalance_status(table)`, reporting skew and whether a rebalance is `none`, `recommended` or `required`. Unlike the preset endpoint, installing centroids does not write postings for existing rows — that path is "install once, then write". Querying an IVF index with no centroids installed returns `503`, not `404`, with the install URLs in the response body. hybrid dense + sparse `POST /vector/:table/topk_hybrid` runs a dense and a sparse query together and fuses the two rankings server-side with Reciprocal Rank Fusion. Fields: `dense_query`, `dense_dim`, `sparse_query_indices`, `sparse_query_values`, `sparse_dim`, `k`, plus optional `dense_metric`, `dense_mode`, `rrf_k` (default 60), `candidates` and `filter`. The returned `score` is a fused rank score, not a distance — it is not comparable to the scores from a single-mode query. No SDK wraps it. 6.7 ## Delete. #### cURL ``` # Single - idempotent, 200 with {"deleted": false} when the id is absent curl -X DELETE \ "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/01JTRX9KQ3YH8K2WMX0F5JZAB7" \ -H "Authorization: Bearer $OC_TOKEN" # Bulk - at most 10,000 ids per call curl -X POST "https://$OC_HOST/v1/tenants/$OC_TENANT/vector/shop.orders/delete-bulk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "ids": ["01JTRX...AB7", "01JTRX...AB8"] }' ``` #### Python ``` db.vector.delete("shop.orders", "01JTRX9KQ3YH8K2WMX0F5JZAB7") out = db.vector.delete_bulk("shop.orders", ["01JTRX...AB7", "01JTRX...AB8"]) print(out.deleted_count, out.missing_count) ``` #### TypeScript ``` const out = await db.vectorDelete("shop.orders", "01JTRX9KQ3YH8K2WMX0F5JZAB7"); console.log(out.deleted); // Bulk delete is not wrapped in the TypeScript SDK - POST /delete-bulk directly. ``` #### Go ``` // The Go SDK has no vector delete method - call the endpoint directly. req, _ := http.NewRequestWithContext(ctx, "DELETE", "https://"+os.Getenv("OC_HOST")+"/v1/tenants/"+os.Getenv("OC_TENANT")+ "/vector/shop.orders/01JTRX9KQ3YH8K2WMX0F5JZAB7", nil) req.Header.Set("Authorization", "Bearer "+os.Getenv("OC_TOKEN")) resp, err := http.DefaultClient.Do(req) ``` - Single delete is idempotent: deleting an id that isn't there is a `200` with `{ "deleted": false }`, never a 404. - Bulk delete caps at 10,000 ids and returns `{ deleted_count, missing_count }`. - Both accept an optional `index` selector (`hnsw` by default) — it must match the index the vector was written under. - Vector ids are capped at 1024 characters and may not be empty. 6.8 ## Limits and gotchas. | Limit | Value | | --- | --- | | Max `k` per query | 4096 | | Max vectors per bulk insert | 100,000 | | Max ids per bulk delete | 10,000 | | Max `nprobe` | 256 | | Max IVF partitions | 65,536 | | Max PQ subspaces (`pq_m`) | 64 | | Vectors sampled for index training | 1,000,000 | | Max vector id length | 1024 chars | | Dimensionality | > 0, no upper bound | metric and index are per-request - and only metric is checked `index` is per-request and recorded nowhere. `metric` is per-request too, but it is checked: every HNSW graph is stamped with the metric it was built under, so writing with `cosine` and querying with `l2` is refused with a `409` `vector_metric_mismatch`. An index built before that stamp shipped carries none and still serves the mismatch behind a `200`, until the next write stamps it. Same for `index`: writing under the default graph index and then querying `index: "ivf_pq"` finds nothing at all, unless the preset build populated those cells. Pick one metric and one index kind per collection and hold them constant in your own code. training and populate sample at most one million vectors A preset build reads up to 1,000,000 vectors, both for training and for populating cells. On a collection larger than that, rows beyond the first million are not written into the IVF-PQ index by the build and will not be found by an `ivf_pq` query until they are written again under that index. back-pressure and quota Queries take a process-wide heavy-operation permit before touching storage; under memory pressure you get `429` with a `Retry-After`. Writes are checked against your vector quota and answer `402` when it is exhausted — a bulk insert is pre-checked for the whole batch, so it is all-or-nothing. Vector endpoints also require the vector capability to be enabled on the instance; without it the call fails with a `402` naming it. Vector search is included on every paid configuration at no extra charge, so this is a switch to flip, not a purchase. quantization on the write path Separately from IVF-PQ, a write can carry `quantization`: `"none"` (default), `"scalar"`, `"binary"` or `"pq"`. This shrinks the stored payload at some cost in precision. Note that under binary quantization `cosine` and `dot` become the same computation, and `manhattan` is evaluated on the `l2` path — rank-equivalent, but the absolute scores differ from what you would expect. 6.9 ## Where this lives in the dashboard. The query workbench has no vector mode. The `LANG` switcher offers SQL, Cypher, NL and Search — and that is the whole list. A nearest-neighbour query needs a query embedding, which is not something you can usefully type into a text editor. What the dashboard does own is the index: on Data → Schema you can train and install a vector index over a table that already holds vectors. Running the search itself is always an API or SDK call. Vector collections do appear in the workbench's schema rail under a vector tables heading with their vector count, so you can confirm what has been indexed without leaving the page. See [Run queries from the dashboard](https://originchaindb.com/docs/dashboard/queries) for the rest. --- # Vector search examples by industry - RAG - OriginChainDB Canonical source: https://originchaindb.com/docs/vector/industries Sitemap last modified: 2026-09-13T20:18:50.000Z vector · by industry # Vector models by industry Four places where matching on meaning beats matching on words. Each one pairs the vector search with the filter or the keyword query that makes it usable. ## Retail: search that understands the ask A shopper types something no keyword index will match, and the catalog still has the right answer. ### Nearest products, within the shopper's constraints The price filter is part of the search, so the ten results you get back are ten results you can show. ``` POST /v1/tenants/:t/vector/shop.products/topk { "query": [...], "k": 10, "dim": 768, "metric": "cosine" } ``` ### Pair it with keywords Vectors are strong on meaning and weak on exact terms - a model number, a brand. Run both and merge; the full-text side is [the search API](https://originchaindb.com/docs/elasticsearch/search). ``` { "query": { "match": { "name": "carbon marathon" } } } ``` ## Support: deflect the ticket before it is filed The customer describes a problem in their words. The answer already exists, written in someone else's. ### Nearest resolved tickets Embed the draft as it is typed and search resolved history; a good match becomes a suggestion instead of a ticket. ``` POST /v1/tenants/:t/vector/support.tickets/topk ``` the two surfaces answer together Nearest-neighbour finds the ones that mean the same thing; [the suggester](https://originchaindb.com/docs/elasticsearch/suggest) fixes a misspelled product name before either search runs. ## Legal and research: retrieval for grounded answers The retrieval half of RAG. The quality ceiling of the answer is set here, not in the model. ### The passages worth sending Chunk the documents, embed the chunks, and retrieve the few that actually bear on the question. ``` POST /v1/tenants/:t/vector/kb.chunks/topk ``` keep the citation Store the source id and offset on the chunk row. Retrieval then returns something a reader can verify, which is the difference between an answer and a claim. [RAG guide](https://originchaindb.com/docs/rag). ## Media: finding the near-duplicates Not identical files, which hashing finds, but the same thing re-encoded, cropped or re-uploaded. ### Neighbours above a similarity threshold Search each new asset against the library and treat anything very close as a candidate duplicate for review. ``` POST /v1/tenants/:t/vector/media.assets/topk ``` ## What these have in common - The vector is a column. It is written with the row, so there is no embedding store to sync and nothing to fall out of date. - Filters belong inside the search. Filtering after a top-k is how you ask for ten results and show three. - Meaning and words are different questions. The strongest results usually come from running both and merging. [Start from zero Declare a column, write a vector, search it.](https://originchaindb.com/docs/vector/quickstart)[Make it fit Quantization and index choice as the corpus grows.](https://originchaindb.com/docs/vector/quantization) --- # IVF vector index on OriginChainDB - nlist, nprobe Canonical source: https://originchaindb.com/docs/vector/ivf Sitemap last modified: 2026-09-13T20:18:50.000Z docs · vector · IVF # IVF vector index. IVF partitions your vectors into cells by their nearest centroid. At query time, the engine picks the `nprobe` closest cells to the query vector and scans only those. The result is the path to 10M+ vectors per tenant on the standard configurations, and a 1M-vector IVF build in 80 s. ## The IVF pipeline. 1. Train — k-means picks `nlist` centroids on a training sample. `sqrt(N)` is a reasonable starting point for `nlist`. 2. Assign — every vector is assigned to its nearest centroid. Bulk-load runs this in parallel: 1M vectors in 80 s. 3. Query — the query vector is scored against the centroids; only the `nprobe` closest cells are scanned. ## IVF is not declared in the manifest. A schema manifest has no `index`, `nlist` or `quantization` key. You choose IVF at runtime, with the [install-centroids call](https://originchaindb.com/docs/vector/ivf#install) below — `nlist` is one of its arguments, not a schema field. Quantization is a separate install call on the same collection, so IVF and [quantization](https://originchaindb.com/docs/vector/quantization) are chosen independently. What the manifest can declare is the collection itself — how wide its vectors are and which distance function they are indexed under. That block is optional, and it is a single `[vector]` table: ``` # manifest.toml — a 768-dim collection. Nothing here selects an index. [vector] dim = 768 # required distance = "cosine" # cosine | l2 | dot | manhattan | l1 ``` Every field, and what happens when a request contradicts the declared distance, is in the [schema reference](https://originchaindb.com/docs/schemas/reference#vector). Registration ignores keys it does not recognise, so an invented one registers with a `200` and silently does nothing — worth knowing before you improvise a field name. ## Install centroids (admin). Training is an HTTP admin call. Run it once when the corpus reaches training scale; re-run only if the data distribution shifts materially. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/vector/embeddings/install-centroids" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "nlist": 1024, "training_set": "sample", "iters": 25 }' ``` ## The nprobe knob. `nprobe` is the recall / latency dial. Higher = better recall, more cells scanned. The recall figures below are the engine's own illustrative curve, measured on a 1,000-vector planted-cluster corpus at 16 partitions — your numbers will move with corpus geometry. nproberecall@10latency band `1` `~0.30` Cheapest scan, lowest p99. `2` `~0.55` Still cheap, recall climbing fast. `4` `~0.85` The default on a 16-partition table. `8` `~0.95` Recall-critical reads. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/vector/embeddings/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [/* 768 floats */], "k": 10, "nprobe": 4 }' ``` ## When to pick IVF over HNSW. - Working set is above ~5M vectors and HNSW memory is starting to bind. - You want to combine with PQ (see [IVF-PQ](https://originchaindb.com/docs/vector/ivf-pq)) for the 64–768× memory story. - Workload is bulk-load heavy. The IVF bulk-load path lands 1M vectors in 80 s. - Recall budget can tolerate the nprobe trade — nprobe=4, the default on a 16-partition table, reaches recall@10 of about 0.85 on that corpus. --- # IVF-PQ on OriginChainDB - residualised product quantization Canonical source: https://originchaindb.com/docs/vector/ivf-pq Sitemap last modified: 2026-09-13T20:18:50.000Z docs · vector · IVF-PQ # IVF-PQ (residualised). IVF-PQ is the residualised version of [IVF](https://originchaindb.com/docs/vector/ivf) from Jegou, Douze, & Schmid (2011). The key correctness point: PQ is not applied to the raw vector — it is applied to the residual of the vector against its assigned centroid. That keeps codebook entries close together in distance space, which keeps recall recoverable even at aggressive memory factors. ## Why residualised matters. A naive PQ over the raw vector loses geometry that the IVF centroid already encodes. Residualised PQ stores only the offset from the assigned centroid; the LUT (look-up-table) at query time is computed per query against the cell centroid, then distances are read off the table in O(M) per candidate. That is the move that makes the memory math work without collapsing recall. ## Memory math. M=16 sub-vectors, ksub=256, residualised per cell. Add a few extra bytes per vector for cell id + tombstone bit. dimraw f32IVF-PQfactor `128` `512 B / vec` `16 B / vec + cell id` `64×` `384` `1.5 KB / vec` `24 B / vec + cell id` `~200×` `768` `3.0 KB / vec` `32 B / vec + cell id` `~400×` `1536` `6.0 KB / vec` `48 B / vec + cell id` `768×` ## Recall floor. The synthetic acceptance floor for IVF-PQ is `recall@10 >= 0.75` at `nprobe=4`. Above that, every CI run for the index path has to clear the same gate or the build fails. Your production numbers will sit above the floor on most real corpuses — Zipfian / clustered distributions are friendlier to IVF than a uniform synthetic. ## The 10M+ scale story. HA+ configurations running IVF-PQ have verified 10M+ vectors in steady state. The IVF bulk-load path lands 1M vectors in 80 s. On Single-zone and HA compute, the practical ceiling is the compute budget rather than the index — IVF-PQ helps but does not make a smaller box infinite. See the [pricing footnote on resources](https://originchaindb.com/pricing) for the honest framing. --- # Vector quantization on OriginChainDB - binary, PQ, IVF-PQ Canonical source: https://originchaindb.com/docs/vector/quantization Sitemap last modified: 2026-09-13T20:18:50.000Z docs · vector · quantization # Vector quantization. Quantization trades a controlled amount of recall for a large reduction in memory. OriginChainDB ships three variants: binary (32× memory), PQ (64× memory at D=128), and IVF-PQ (residualised per-cell, 768× memory at D=1536). All three are GA. Pick one per collection at install time; the topk surface stays the same. ## The trade table. kindmemoryrecall bandwhen to pick `Scalar (f32)` `1× (baseline)` Reference - the recall every other variant trades against. Small indexes, recall-critical workloads, anything under ~1M vectors where memory isn't the binding constraint. `Binary quantization` `32×` recall@10 typically 0.85–0.93 vs f32; rerank top-100 to recover. Memory-bound workloads. Pair with HNSW or IVF. Cheapest scoring kernel - hamming distance on packed bits. `PQ (8-bit codebook)` `64× at D=128` recall@10 typically 0.92–0.96 with M=16 sub-vectors. When you want compressed storage but don't need IVF partitioning. LUT scoring at query time keeps latency similar to f32. `IVF-PQ (residualised)` `64× at D=128, 768× at D=1536` CI acceptance floor: recall@10 >= 0.75 at nprobe=4 on a small planted-cluster corpus. 10M+ vectors per tenant. IVF partitions + per-cell PQ codebook (Jegou 2011). ## Pick one at install. Quantization is a per-collection property picked at install time, not in the schema. There is no `quantization` key in a manifest — you choose it with an admin call against the collection you are already writing to, the way [IVF centroids](https://originchaindb.com/docs/vector/ivf#install) are installed. The topk endpoint does not change; the engine picks the scoring kernel for you. It is independent of the index kind, so pair binary or PQ with HNSW for default recall, or with [IVF](https://originchaindb.com/docs/vector/ivf) for larger corpuses. The manifest's part in this is only the optional `[vector]` table, which declares the collection's `dim` and `distance` — see the [schema reference](https://originchaindb.com/docs/schemas/reference#vector). Registration ignores keys it does not recognise, so a made-up quantization field would register with a `200` and quietly do nothing. ## Train a PQ codebook. PQ needs a codebook trained on a sample of your data. The training set is a row label that the engine sub-samples; sub-vectors (`m`) and codebook entries per sub-vector (`ksub`) are the two knobs. Defaults are `m=16, ksub=256` for D=128. #### cURL ``` curl -X POST "https://acme.db.originchain.ai/v1/tenants/$T/vector/embeddings/install-pq" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "m": 16, "ksub": 256, "training_set": "sample" }' ``` ## Memory math. At D=128 with f32, one vector costs 512 bytes. Binary collapses that to 16 bytes (1 bit × 128 dims). PQ with M=16 sub-vectors stores 16 codes per vector at 8 bits each = 16 bytes. IVF-PQ residualises the per-cell vector and stores 16 bytes plus the cell id; at D=1536 the residual structure lands the 768× figure. --- # Vector search quickstart - top-k and filters - OriginChainDB Canonical source: https://originchaindb.com/docs/vector/quickstart Sitemap last modified: 2026-09-23T08:33:23.000Z vector · quickstart # Run your first vector search Store an embedding and run your first nearest-neighbor query with OriginChainDB. ## Before you start An instance, an API key, and a model that turns your content into vectors. The dimension you choose is fixed per column. ``` export OC_URL='https://' export OC_TENANT='' export OC_TOKEN='' ``` ## The four steps 1. 1. Declare a vector column The type is spelled "vector". Its width goes in the column's [columns.params] block as vector_dim, and the optional top-level [vector] table pins the collection's dimension and distance metric. Declare both and the two dimensions must agree, or registration fails — see the [schema reference](https://originchaindb.com/docs/schemas/reference#vector) and [vector schema](https://originchaindb.com/docs/schemas/vector). ``` [[columns]] name = "embedding" ty = "vector" [columns.params] vector_dim = 768 # Optional collection config. If you declare it, its dim must # equal the column's vector_dim or registration fails. [vector] dim = 768 distance = "cosine" ``` 2. 2. Write the row, then the vector Vector values are not stored in the row body. They live in their own keyspace, keyed by (tenant, table, id), and are written through the vector-put path — sending an embedding to the rows endpoint does not index it. Reuse the row's id on both sides and a search hit reads straight back to the full row. No second system to keep in step. ``` // 1. the row - every column except the embedding curl -X POST "$OC_URL/v1/tenants/$OC_TENANT/rows/shop.products" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"id":"sku-8842","name":"Carbon Marathon","price":149}' // 2. the embedding - same id, vector keyspace curl -X POST "$OC_URL/v1/tenants/$OC_TENANT/vector/shop.products/put" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{"id":"sku-8842","dim":768, "embedding":[0.011,-0.082,0.046]}' // 768 floats in practice ``` 3. 3. Search for the nearest Send a query vector and how many neighbours you want. ``` curl -X POST "$OC_URL/v1/tenants/$OC_TENANT/vector/shop.products/topk" \ -H "Authorization: Bearer $OC_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": [0.011, -0.082, 0.046], "k": 10, "dim": 768, "metric": "cosine" }' ``` 4. 4. Filter while you search Metadata filters apply during the search, not after it, so asking for ten results under a price cap returns ten — not whatever survives a post-filter. ``` // see the reference for the filter syntax and how it interacts with recall ``` pick the metric to match the model Use the metric your embedding model was trained for. cosine is the usual answer for text embeddings; the wrong metric quietly returns plausible, worse neighbours rather than failing. ## What to read next [Vector reference Top-k, filters, metrics, and the speed-recall trade.](https://originchaindb.com/docs/vector)[IVF and IVF-PQ Index choices for scale, and what each costs in recall.](https://originchaindb.com/docs/vector/ivf)[By industry Semantic product search, support deflection, document retrieval.](https://originchaindb.com/docs/vector/industries) --- # Data Processing Addendum - OriginChain Canonical source: https://originchaindb.com/dpa Sitemap last modified: 2026-09-23T08:33:23.000Z legal # Data Processing Addendum Last updated: 4 September 2026 Browse this document - [1. Parties & scope](https://originchaindb.com/dpa#dpa-section-1) - [2. Roles](https://originchaindb.com/dpa#dpa-section-2) - [3. Categories of data & data subjects](https://originchaindb.com/dpa#dpa-section-3) - [4. Sub-processors](https://originchaindb.com/dpa#dpa-section-4) - [5. International transfers](https://originchaindb.com/dpa#dpa-section-5) - [6. Security](https://originchaindb.com/dpa#dpa-section-6) - [7. Data subject requests](https://originchaindb.com/dpa#dpa-section-7) - [8. Deletion & retention](https://originchaindb.com/dpa#dpa-section-8) - [9. Breach notification](https://originchaindb.com/dpa#dpa-section-9) - [10. Audit](https://originchaindb.com/dpa#dpa-section-10) - [11. Liability](https://originchaindb.com/dpa#dpa-section-11) - [12. Governing law](https://originchaindb.com/dpa#dpa-section-12) - [13. Contact](https://originchaindb.com/dpa#dpa-section-13) ## 1. Parties & scope This Addendum forms part of the Master Services Agreement between the customer ("Controller") and Silicoyn Technologies Pvt Ltd ("OriginChain" or "Processor"). It governs the processing of Personal Data the Controller submits to the OriginChain Service. ## 2. Roles The Controller is the data controller; OriginChain is the data processor. OriginChain processes Personal Data only on documented instructions from the Controller - typically the API and console actions taken by the Controller's authenticated users. ## 3. Categories of data & data subjects The Service is a general-purpose database; the Controller decides what to store. OriginChain does not classify Personal Data on the Controller's behalf. Typical categories include: - End-user identifiers (email, account id) - Application content (rows, schema, query history) - Operational telemetry (request timestamps, error codes) Special-category data (health, biometrics, political opinion etc.) may only be processed under an Enterprise contract with appropriate sector-specific addenda. ## 4. Sub-processors OriginChain uses the following sub-processors as of the date above. We notify Controllers in advance of any addition or replacement. - Amazon Web Services Cloud infrastructure: compute, storage, backups, DNS, object storage, and our managed account database. - Razorpay Payment processing. Card data is tokenised - OriginChain never receives PAN or CVV. - Model provider (Controller-selected) Natural-language compilation of `/v1/ask` requests. The Controller supplies its own model-provider credentials, so the provider is whichever the Controller has configured. Prompts are submitted to that provider and discarded; OriginChain does not store model output beyond the response it returns. ## 5. International transfers Customer data is stored in the region selected by the Controller and does not leave it. Live regions span Asia-Pacific, Europe, North and South America, Africa and the Middle East; the current list is published by the control plane and shown in the console when an instance is created. A Controller that must keep data within a specific jurisdiction should select a region in that jurisdiction. Cross-border transfers, if any, occur only when the Controller routes AI traffic through a model endpoint in another region - controlled by the Controller's tenant configuration. ## 6. Security See [/security](https://originchaindb.com/security) for the full posture. Summary: TLS for every external endpoint, Argon2id for stored credentials, isolated tenant resources, daily backups encrypted at rest, and principle-of-least-privilege engineer access that is brokered and audit-logged (no direct SSH). ## 7. Data subject requests On the Controller's instruction, OriginChain will assist with the rectification, erasure, restriction, or export of Personal Data. Self-serve mechanisms: - Erasure: cancel a subscription with the "delete data" option, or contact [privacy@originchain.ai](mailto:privacy@originchain.ai) to expedite. - Export: scan the affected schemas via the API. We do not gate exports. - Rectification: standard write API. ## 8. Deletion & retention On termination of a tenant subscription with the "delete data" flag set: - The per-tenant compute and storage volume are destroyed within 24 hours. - Backup snapshots age out after 30 days per the standard retention policy. Earlier deletion is available on written request. - Operational metadata in our management systems (instance and subscription rows) is retained for 7 years to satisfy financial record-keeping obligations under Indian tax law (Section 44AA, Income Tax Act, 1961). - Security logs (auth events, failed login attempts) are retained for 90 days. ## 9. Breach notification OriginChain will notify the Controller without undue delay, and in any event within 72 hours of becoming aware, of any Personal Data Breach affecting the Controller's data. Notification includes the nature of the breach, the categories and approximate number of records affected, the likely consequences, and the measures taken or proposed to address it. ## 10. Audit OriginChain provides on request: SOC 2 Type 1 reports as they are issued (Type II available on request once attested), the security questionnaire (today: [/security](https://originchaindb.com/security)), and access to engagement-specific evidence (penetration test reports, vulnerability scans, access logs filtered to the Controller's tenant). On-site audits are negotiated per Enterprise contract. ## 11. Liability Liability under this Addendum is governed by the limitation-of-liability clause in the Master Services Agreement. ## 12. Governing law This Addendum is governed by the laws of India. Disputes are subject to the exclusive jurisdiction of the courts in Bengaluru, Karnataka. Where a Controller is established in the EU/UK and EU/UK data protection law applies to the processing, the parties incorporate the EU Standard Contractual Clauses (Module Two: Controller to Processor) as a separate signed annex. ## 13. Contact Privacy office: [privacy@originchain.ai](mailto:privacy@originchain.ai) Postal address: Silicoyn Technologies Pvt Ltd, Bengaluru, Karnataka, India. --- # Dedicated database deployments | OriginChainDB Canonical source: https://originchaindb.com/enterprise Sitemap last modified: 2026-09-23T08:33:23.000Z enterprise # Dedicated database deployments Run OriginChainDB on dedicated infrastructure in your chosen region, with deployment, recovery and support requirements agreed for your workload. [Book a technical walkthrough](https://originchaindb.com/contact?topic=walkthrough) [Start free](https://app.originchaindb.com/signup) 99.99% 1 monthly availability · Enterprise 35 days 2 recovery window, up to 4,940 ops/s 3 measured throughput · YCSB 1. 1 Contractual — Enterprise agreement. /terms § Availability — monthly uptime on Enterprise. 2. 2 Control plane `MAX_PITR_RETENTION_DAYS = 35` (operation-backend src/instances/handlers.rs); the console offers 7- and 30-day windows. 3. 3 Measured — YCSB, client fan-out; hardware and client count in the methodology section. /benchmarks § 01 YCSB — 50,000 ops across workloads A–F on a 20K-row table, zero errors. 01 · what you get ## One engine. Every guarantee written down. Every figure below is drawn from the canonical claims sheet that also feeds our pricing page, FAQ and terms — the same figure appears everywhere, or it appears nowhere. ### Dedicated, in your region Every dedicated configuration is single-tenant, region-isolated: its own compute, storage, hostname and bearer token. Data, transaction logs, backups and the natural-language compile step stay inside the region you choose. ### Distributed by design Active-passive replication is the shipped path and the default: one node holds the writer lease while provisioned standbys tail its stream and can be promoted. Promotion is operator-run by default. An opt-in automatic path can promote a standby once the writer's lease has expired — it stays off unless enabled, relies on an external restart hook, and refuses to promote a standby that is not fully caught up. Quorum-commit replication is roadmap and not scheduled: a request to run a database that way is refused rather than accepted and degraded, and a running database cannot be switched into a quorum group today. ### Contractual availability 99.9% on HA, 99.95% on HA+ and 99.99% on Enterprise, with a pro-rata service credit against the next invoice. Summarised on /sla; /terms governs. ### Recovery you can put in a runbook encrypted nightly snapshots, included. Opt-in point-in-time recovery restores to a roll-up boundary, and the spacing of those boundaries follows a configured shipping interval that the Intra-Segment PITR add-on tightens. Retention up to 35 days. ### Transactions when you need them single-row atomic writes with optimistic concurrency (If-Match / _oc_row_version), included. Add multi-row transactions under snapshot isolation with the Transactions add-on. ### Security controls on every configuration TLS 1.2+ in transit · AES-256 at rest · customer-managed keys on Enterprise. per-request audit log on every configuration. HIPAA BAA and GDPR DPA are available on Enterprise agreements. 02 · procurement ## Built to pass the review. Security, legal and finance each get a document they can file. Anything not listed here is scoped in the first call. [Request the security pack](https://originchaindb.com/contact?topic=enterprise) [Read the trust center](https://originchaindb.com/trust) Agreement Annual contract with negotiated quotas, custom retention and a named support engineer. Availability 99.99% monthly uptime target on Enterprise; credits per /terms. Data protection GDPR DPA and HIPAA BAA on request; sub-processor details are provided with the DPA. Security review Security questionnaire, architecture walkthrough and current SOC 2 letter (audit in progress) on request. Keys Customer-managed encryption keys on Enterprise. Deployment Managed in-region by default; private deployments are scoped per engagement. 03 · how enterprises start ## A pilot with success criteria, not a demo. We stand up a dedicated instance against your workload shape, agree the numbers that matter before the first byte lands, and hand you a written report at the end. The structure is on the [pilot page](https://originchaindb.com/pilot). ## Talk to an engineer, not a deck. A technical walkthrough is a working session on your workload: data shapes, query patterns, availability target, region. You leave with a sized configuration and a written quote. [Book a technical walkthrough](https://originchaindb.com/contact?topic=walkthrough) [Start free](https://app.originchaindb.com/signup) [For developers Follow the database quickstart](https://originchaindb.com/docs/quickstart) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB events and sessions Canonical source: https://originchaindb.com/events Sitemap last modified: 2026-09-23T08:33:23.000Z Events & sessions # Meet the ideas. Build on the details. Conversations about data models, queries and the applications they make possible. 01 / The schedule ## Upcoming ### Our next session is not announced yet. Planning a conference, meetup or team workshop? Tell us about the audience and what you would like to explore. [Discuss a session](https://originchaindb.com/contact) 02 / The archive ## From the stage. Revisit the session details and available materials from our past events. 1. Past eventConference · Speaking ### Global Fintech Fest 2026 A session on the data layer for enterprise AI and agentic workflows. When 8–11 Sep 2026 Where Mumbai, IndiaJio World Centre, BKC The session #### One AI multimodal database platform for enterprise data, AI and agentic workflows Zaheer Kazi, Co-founder & Chief Commercial Officer The discussion connected structured records, document search, embeddings and entity relationships to the context an AI application needs. [[Image: Zaheer Kazi speaking on the Global Fintech Fest main stage, with the OriginChainDB opening slide on the screen behind him.]](https://originchaindb.com/events/gff-2026/opening.webp) Opening the session on the main stage. More from the session 4 photos [[Image: Wide view of the stage during the session, the audience-facing screen showing a slide about the data layer for agentic AI.]](https://originchaindb.com/events/gff-2026/data-tier.webp) Why the data layer is where AI projects stall. [[Image: Closer view of the same slide on stage, with the speaker at the right of the frame.]](https://originchaindb.com/events/gff-2026/data-tier-close.webp) [[Image: The speaker presenting a slide listing the ways teams adopt the database, from individual builders to enterprise programmes.]](https://originchaindb.com/events/gff-2026/go-to-market.webp) [[Image: The closing slide of the session, reading 'One source of truth. Every query shape.']](https://originchaindb.com/events/gff-2026/closing.webp) One source of truth. Every query shape. ## Take the next step. Choose the path that fits what you want to learn. [Understand ### Explore the database How SQL, vector, graph and full-text work with your data.](https://originchaindb.com/product/database) [Build ### Run a first query Go from the product overview to the API and query examples.](https://originchaindb.com/docs/quickstart) [Discuss ### Join the community Bring your implementation questions to the conversation.](https://community.originchain.ai/) --- # OriginChainDB questions and answers Canonical source: https://originchaindb.com/faq Sitemap last modified: 2026-09-23T08:33:23.000Z OriginChainDB / Questions & answers # Find the answer. Follow the details. Data models, deployment, pricing and the limits to understand before you build. [Implementation guides](https://originchaindb.com/docs)[Ask the team](https://originchaindb.com/contact) ### Need a specific answer? Bring your workload or implementation question to the team. [support@originchain.ai](mailto:support@originchain.ai) ## Getting started What OriginChainDB is, what you get, and how to pick a starting point. ### What is OriginChainDB? OriginChainDB is a distributed multimodal database for SQL records, vector search, full-text search and graph relationships. Ask adds a natural-language query interface over the engine. [Explore the database](https://originchaindb.com/product/database). [Link to this answer](https://originchaindb.com/faq#what) ### What do I get when I sign up? A provisioned database has an HTTPS endpoint and credentials for its supported API surfaces. Free uses shared infrastructure; dedicated configurations have their own compute and storage. Review the selected configuration and [follow the quickstart](https://originchaindb.com/docs/quickstart) for your first request. [Link to this answer](https://originchaindb.com/faq#what-you-get) ### How much does OriginChainDB cost? Free includes 1 GB of database storage. Two pooled configurations are available: $29/month for 25 GB storage and 20 GB monthly data transfer, or $49/month for 50 GB storage and 50 GB transfer. Dedicated pricing depends on resources, region and options. [Compare configurations and estimates](https://originchaindb.com/pricing). [Link to this answer](https://originchaindb.com/faq#pricing) ### Which regions can I deploy to? Check the regions available for the configuration you select in the console. Dedicated database resources are provisioned in the selected region; account and billing services are global. The pooled configurations are scoped to India (Mumbai). [Read the deployment guide](https://originchaindb.com/docs/deploy). [Link to this answer](https://originchaindb.com/faq#regions) ### How do I sign up and get my first connection? Create an account at [app.originchaindb.com](https://app.originchaindb.com/signup), select an available configuration and review its requirements. Once provisioned, use the instance endpoint and credential from the console. Keep credentials on your server and follow the [first-query guide](https://originchaindb.com/docs/quickstart). [Link to this answer](https://originchaindb.com/faq#signup) ### Why is the dashboard on a different domain? The public website at originchaindb.com and the authenticated console at app.originchaindb.com are separate applications. Use the website for product information and documentation, and the console for your account and instances. [Link to this answer](https://originchaindb.com/faq#dashboard-host) ### Free, Starter or dedicated — which should I pick? Use Free for learning and prototypes. The 25 GB and 50 GB pooled configurations provide fixed monthly shared capacity; choose your configuration in the console. Choose a dedicated configuration when you need your own compute and storage, with recovery and standby options reviewed separately. [Compare the current options](https://originchaindb.com/pricing). [Link to this answer](https://originchaindb.com/faq#which-plan) ## Capabilities Data models, query interfaces, transactions and supported operations. ### What query shapes does OriginChainDB support? The four data models are SQL records, vectors, full-text indexes and graph relationships. Ask is a natural-language interface, not a fifth storage model. Each API has its own supported operations and limits; the [documentation](https://originchaindb.com/docs) shows the request and result shapes. [Link to this answer](https://originchaindb.com/faq#query-shapes) ### What are 'atomic cross-shape writes'? Atomicity applies to the writes included in a supported operation or transaction. A row update does not automatically update a separately written vector or full-text index. Coordinate those index writes in your application and use stable record IDs to connect results. See [transaction guarantees](https://originchaindb.com/docs/schemas/transactions) and [full-text indexing](https://originchaindb.com/docs/fts). [Link to this answer](https://originchaindb.com/faq#atomic-cross-shape) ### Is OriginChainDB a distributed database? Dedicated configurations can use a primary and optional standby with asynchronous replication. The primary accepts writes; promotion and availability depend on the configured topology. This is not a promise of quorum commit or multiple active writers. See [replication and recovery behavior](https://originchaindb.com/docs/ops) and your [service agreement](https://originchaindb.com/sla). [Link to this answer](https://originchaindb.com/faq#distributed) ### Can vector search and SQL hit the same row? Yes. Use the same stable ID for a SQL row and its vector entry, then fetch the row for a returned search ID. The vector is written through its own API; changing the row does not automatically regenerate its embedding. [See how vector entries relate to rows](https://originchaindb.com/docs/schemas/vector). [Link to this answer](https://originchaindb.com/faq#mix-shapes) ### How does natural-language /ask work? Ask first tries a deterministic grammar against the schema catalog. If that cannot resolve the question, it uses the configured LLM; a compiled plan can then be cached. Pass a schema scope and use show_plan to inspect the result. Schema changes can invalidate cached plans. [Read the Ask reference](https://originchaindb.com/docs/ask). [Link to this answer](https://originchaindb.com/faq#natural-language) ### Do I bring my own LLM key? Ask can fall back to the LLM configured for the instance. Review the provider, credential handling and provider terms before enabling that path. The deterministic grammar and plan cache can answer supported requests without a new LLM call. [See how Ask chooses its path](https://originchaindb.com/docs/ask). [Link to this answer](https://originchaindb.com/faq#byok-llm) ### What does the AI compiler see? An Ask question is sent to the configured LLM when the deterministic path needs that fallback. Do not include secrets or sensitive personal data in the question. Review the configured provider and data handling requirements rather than assuming every external inference call stays in the database region. [Read the request guidance](https://originchaindb.com/docs/ask). [Link to this answer](https://originchaindb.com/faq#ai-data) ### Does OriginChainDB support graph queries? Yes. Declare relations on registered tables, then follow connected records or use supported graph algorithms. Cypher and algorithm endpoints have different syntax, limits and result shapes; supported Cypher is a subset. [Review traversal examples and limits](https://originchaindb.com/docs/graph). [Link to this answer](https://originchaindb.com/faq#graph) ### What languages does full-text search support? The HTTP API uses Unicode tokenization and lowercase normalization across scripts. Engine stemmers and other analyzer stages are not all selectable through the API: do not assume stemming, lemmatization or accent folding is enabled. [Check the live analyzer behavior and synonym support](https://originchaindb.com/docs/fts). [Link to this answer](https://originchaindb.com/faq#fts-languages) ### Does OriginChainDB support transactions? Single-row writes are atomic. The preview /tx API defaults to snapshot isolation with commit-time conflict detection and an optional serializable mode; SQL session transactions use read committed with different conflict behavior. Pick the surface for your workload and handle commit failures explicitly. [Compare the two transaction APIs](https://originchaindb.com/docs/schemas/transactions). [Link to this answer](https://originchaindb.com/faq#transactions) ## Performance How to evaluate latency, throughput, recall and dataset size. ### What latencies should I expect? Latency depends on query shape, data size, indexes, hardware, load and network distance. Measure the full request path on your own dataset; a published benchmark is evidence for its stated setup, not a latency guarantee for every configuration. [Review the benchmark methods and results](https://originchaindb.com/benchmarks). [Link to this answer](https://originchaindb.com/faq#latency) ### How fast is /ask? A cache hit, a fresh grammar compilation and an LLM fallback have different costs. Inspect the cache field and query plan, then measure the path your application uses. There is no single response time that covers all three. [See plan-cache behavior](https://originchaindb.com/docs/ask). [Link to this answer](https://originchaindb.com/faq#ask-latency) ### How accurate and fast is vector search? Vector accuracy and latency depend on the index, dimensions, metric, filters and search settings. Compare recall and tail latency together using a representative query set; fast and high_recall modes make different tradeoffs. [Read published measurements](https://originchaindb.com/benchmarks) and [choose search settings](https://originchaindb.com/docs/vector). [Link to this answer](https://originchaindb.com/faq#vector-perf) ### What throughput can I expect? Use the configuration estimator for capacity planning, then benchmark your workload with representative concurrency, writes and queries. Sizing estimates are not measured throughput guarantees, and adding a standby does not imply linear query scaling. [Review the benchmark resources](https://originchaindb.com/benchmarks). [Link to this answer](https://originchaindb.com/faq#throughput) ### How many vectors can I store? Capacity depends on vector dimensions, metadata, the index representation and available memory and storage. A documented index option does not establish measured performance at every corpus size. [Review the index choices](https://originchaindb.com/docs/vector) and [discuss a representative capacity test](https://originchaindb.com/contact). [Link to this answer](https://originchaindb.com/faq#vector-scale) ## Operations Backups, recovery, failover, observability and SDKs. ### How are backups handled? Dedicated configurations include a daily storage snapshot. Snapshot recovery and continuous archive are separate paths; the operations guide documents current archive limitations. Free and pooled configurations do not include point-in-time recovery. [Review backups and recovery expectations](https://originchaindb.com/docs/ops). [Link to this answer](https://originchaindb.com/faq#backups) ### Do you support point-in-time recovery? Continuous archive and restore-to-timestamp options for dedicated configurations are in preview. The operations guide records lagging or stalled archive progress, so do not treat sub-second or per-second recovery as delivered behavior. Plan against the available daily snapshot and review your recovery requirements with the team. [Read the current recovery caveats](https://originchaindb.com/docs/ops). [Link to this answer](https://originchaindb.com/faq#pitr) ### How does failover work? A configured standby follows the primary asynchronously, so abrupt primary loss can lose recent acknowledged writes. Promotion is fenced by the writer lease; automatic promotion is opt-in and depends on the configured path and catch-up state. A standby does not replace a backup. [Read the failover procedure and durability caveats](https://originchaindb.com/docs/ops). [Link to this answer](https://originchaindb.com/faq#failover) ### What observability do I get? Health and readiness endpoints expose instance status; the usage endpoint reports per-tenant counts and activity. SQL EXPLAIN helps inspect query plans, and the operations guide describes tracing and monitoring options. Confirm the signals available on your deployed configuration. [Explore operational monitoring](https://originchaindb.com/docs/ops). [Link to this answer](https://originchaindb.com/faq#observability) ### Which SDKs are available? Python, TypeScript and Go clients cover different parts of the HTTP API. Python has the widest coverage; TypeScript also offers a control-plane client, while Go covers query surfaces. Use plain HTTP for operations your chosen SDK does not wrap. [Compare the client coverage](https://originchaindb.com/docs/sdk). [Link to this answer](https://originchaindb.com/faq#sdks) ### Can I subscribe to live updates? The HTTP reference documents a watch interface for change notifications. Check its supported subscription scope, delivery behavior and client handling before using it to keep another system in sync. Do not assume it is a durable replay log or a universal CDC replacement. [Read the HTTP API reference](https://originchaindb.com/docs/api). [Link to this answer](https://originchaindb.com/faq#watch) ## Security & compliance Tenancy, encryption, audit logging, residency and certifications. ### Is OriginChainDB single-tenant? Dedicated configurations have their own compute and storage. Free and the $29 and $49 pooled configurations use shared infrastructure. Database isolation, network paths and access controls should be evaluated for the chosen deployment. [Review the deployment boundaries](https://originchaindb.com/security). [Link to this answer](https://originchaindb.com/faq#single-tenant) ### How is my data encrypted? Managed public endpoints use HTTPS. The security page describes where TLS terminates and how the network path differs between shared and dedicated deployments. Review storage, backup and key-management requirements for your selected configuration. [Read the security overview](https://originchaindb.com/security). [Link to this answer](https://originchaindb.com/faq#encryption) ### Is there an audit log? Review which identity, request and outcome fields your deployed service records, where you can access them, and the retention or export options in your agreement. Console preview screens alone do not establish live audit coverage. [Review access and operational records](https://originchaindb.com/security). [Link to this answer](https://originchaindb.com/faq#audit-log) ### What about SOC 2, HIPAA, and GDPR? Contact [security@originchain.ai](mailto:security@originchain.ai) for current assessment status, questionnaires and contractual requirements. A database feature is not a compliance certification; evaluate the deployment and applicable agreement for your workload. [Review the published security information](https://originchaindb.com/security). [Link to this answer](https://originchaindb.com/faq#compliance) ### Where exactly does my data live? Dedicated database resources are provisioned in the selected region. Account and billing are global services; external LLM providers and other integrations have their own data paths. Confirm database, backup and integration requirements together for your deployment. [Read the deployment guide](https://originchaindb.com/docs/deploy). [Link to this answer](https://originchaindb.com/faq#residency) ## Billing What things cost, what is included, and how changes are charged. ### How do add-ons work? The dedicated estimator includes the four data models without separate model fees. Resource sizing, standby, recovery, networking and other agreed options affect the configuration. Confirm the available capabilities and final quote before launch. [Review the configuration options](https://originchaindb.com/pricing). [Link to this answer](https://originchaindb.com/faq#addon-model) ### What happens if I upgrade or add an add-on mid-cycle? Commercial timing follows the applicable agreement; the published Terms distinguish immediate upgrades from next-cycle downgrades. Do not assume all changes are available in place or have no interruption. Free, pooled and dedicated configurations currently do not support in-place switching between them. [Review the Terms](https://originchaindb.com/terms) and [discuss a change](https://originchaindb.com/contact?topic=pricing). [Link to this answer](https://originchaindb.com/faq#mid-cycle) ### Is there an annual discount? Contact the team to agree a billing term or custom commitment. The public estimator shows a monthly equivalent and does not apply an automatic annual discount. [Discuss commercial terms](https://originchaindb.com/contact?topic=enterprise). [Link to this answer](https://originchaindb.com/faq#annual) ### What if I go over my configuration's quota? The $29 configuration includes 25 GB storage and 20 GB monthly transfer; the $49 configuration includes 50 GB storage and 50 GB monthly transfer. Check current usage limits and billing terms in the console before purchasing. For other configurations, review included usage and the final agreement. [See current allowances](https://originchaindb.com/pricing). [Link to this answer](https://originchaindb.com/faq#overage) ### What's included in Enterprise? Enterprise is scoped around your deployment, network, region, recovery, support and contractual requirements. Confirm each capability and service commitment with the team rather than assuming a preset package or an unshipped topology. [Discuss a custom deployment](https://originchaindb.com/contact?topic=enterprise). [Link to this answer](https://originchaindb.com/faq#enterprise) ### Do vector, graph and full-text cost extra? The dedicated pricing estimator includes SQL, vector, graph and full-text without separate per-model fees. Rows and indexes still consume compute, memory and storage, so workload size affects the estimate. [Size a configuration](https://originchaindb.com/pricing). [Link to this answer](https://originchaindb.com/faq#modality-fees) ### What is the Starter plan? The current shared options are Configuration 25 GB pooled at $29/month with 20 GB monthly transfer, and Configuration 50 GB pooled at $49/month with 50 GB transfer. Both are monthly USD configurations with shared compute. Select a configuration in the console and review checkout before paying. [See the current configuration names and limits](https://originchaindb.com/pricing). [Link to this answer](https://originchaindb.com/faq#starter) ### What does the cheapest dedicated configuration include? The public estimator starts at $320/month for a Mumbai baseline with 2 vCPU, 8 GB RAM and 20 GB storage. This is a planning estimate in USD before applicable taxes; availability and the final quote are confirmed before launch. [Review and adjust the estimate](https://originchaindb.com/pricing). [Link to this answer](https://originchaindb.com/faq#cheapest-dedicated) ### What happens when I outgrow Starter? Review a dedicated configuration when you need more capacity or dedicated resources. In-place switching from pooled to dedicated is not currently available; plan the data movement and application cutover with the team. [Discuss the deployment path](https://originchaindb.com/contact?topic=pricing). [Link to this answer](https://originchaindb.com/faq#outgrow-starter) Put the answers to work ## Start with your own workload. Explore the API, compare configurations, or talk through a deployment with the team. [Run a first query](https://originchaindb.com/docs/quickstart)[Compare configurations](https://originchaindb.com/pricing) --- # Database workloads by industry | OriginChainDB Canonical source: https://originchaindb.com/industries Sitemap last modified: 2026-09-23T08:33:23.000Z Industry workflows # Your industry. Four ways to query. Build with SQL, vector, graph and full-text in one database—from financial investigations to product discovery. [Find your industry](https://originchaindb.com/industries#industry-directory)[Meet the database](https://originchaindb.com/product/database) OriginChainDB / Distributed multimodal database Four models, in practiceIllustrative workflows Your applicationAccounts, case notes and policies SQLFilter account activity VectorFind similar case notes GraphFollow account relationships Full-textSearch policy wording Bring the results together Give an investigator records, relationships and supporting evidence. Your application supplies embeddings, indexes text separately and connects results by record ID. ## Start with your workflow. 13 industry guides. Practical places to begin. 01 / 8 industries ### Financial services Put transactions, relationships and documents in context. [#### Banking Connect account activity, customer relationships and case evidence. Explore the workflow](https://originchaindb.com/industries/banking)[#### Payments Investigate transaction patterns alongside merchant and account context. Explore the workflow](https://originchaindb.com/industries/financial-services-payments)[#### NBFCs Bring borrower records and supporting loan documents together. Explore the workflow](https://originchaindb.com/industries/nbfcs)[#### Lending Retrieve application data and the evidence behind a credit review. Explore the workflow](https://originchaindb.com/industries/lending)[#### Insurance Explore claims, related entities and relevant policy wording. Explore the workflow](https://originchaindb.com/industries/insurance)[#### Wealth management Find research alongside client and portfolio context. Explore the workflow](https://originchaindb.com/industries/wealth-management)[#### Mutual funds Connect fund information, investor records and distribution relationships. Explore the workflow](https://originchaindb.com/industries/mutual-funds)[#### Capital markets Explore market records, connected entities and research documents. Explore the workflow](https://originchaindb.com/industries/capital-markets) 02 / 2 industries ### Customer businesses Help people find the right product or the next useful answer. [#### E-commerce & retail Build product discovery with catalog filters, similarity and keyword search. Explore the workflow](https://originchaindb.com/industries/ecommerce-retail)[#### Contact centres Bring customer history and relevant knowledge into support workflows. Explore the workflow](https://originchaindb.com/industries/bpos-contact-centres) 03 / 3 industries ### Essential services Follow operational connections and find the source information. [#### Healthcare Find documents and connected records for care and service workflows. Explore the workflow](https://originchaindb.com/industries/healthcare)[#### Telecommunications Navigate customer, service and network relationships. Explore the workflow](https://originchaindb.com/industries/telecom)[#### Airlines & aviation Connect passenger, route and operations data with searchable knowledge. Explore the workflow](https://originchaindb.com/industries/airlines-aviation) From a use case to a working example ## Bring your data. Build the workflow. Start with a record and add the query models your application needs. [For developersBuild your first example](https://originchaindb.com/docs/quickstart)[For evaluating teamsDiscuss your requirements](https://originchaindb.com/contact) --- # Multimodal data for aviation | OriginChainDB Canonical source: https://originchaindb.com/industries/airlines-aviation Sitemap last modified: 2026-09-23T08:33:23.000Z [← All industries](https://originchaindb.com/industries) Airlines & aviation # One flight changes. Follow the connections. Bring schedules, aircraft rotations, passenger journeys and source documents into an operations application. [Start with a relationship](https://originchaindb.com/docs/graph/quickstart)[Discuss your workflow](https://originchaindb.com/contact) ## An arrival is part of a bigger picture. Start from the flight record, then explore the relationships or documents relevant to the task. Flight context explorerIllustrative application data Starting recordFlight 207`flight:207` City ACity B Late arrival Follow the context ### Which later legs share this aircraft? AircraftAssigned aircraft`aircraft:12` RotationLinked flight legs`rotation:07` Next flightNext scheduled leg`flight:318` Query choicesGraph + SQL Traverse the rotation, then load schedule and crew-assignment records for the connected legs. [Build a relationship query](https://originchaindb.com/docs/graph/quickstart) Illustrative records. Your application assembles this context for the operations team; the team chooses the next action. Three places to begin ## Build around the team's next question. 01 ### Disruption investigation Find affected rotations and onward itineraries, then load schedules and assignment records for review. [Explore relationships](https://originchaindb.com/docs/graph/quickstart) 02 ### Operations knowledge Retrieve similar case notes and exact procedure references, retaining the source document and revision. [Build knowledge retrieval](https://originchaindb.com/solutions/rag) 03 ### Passenger service context Bring booking details, connected journeys and prior service cases into an agent's workspace. [Explore service workflows](https://originchaindb.com/industries/bpos-contact-centres) Four models, specific roles ## Keep the record. Add the context. Use stable flight, aircraft, itinerary and document IDs to connect representations in your application. [Inside the database](https://originchaindb.com/product/database) SQL Filter schedules, bookings and service records by their fields.[Query records](https://originchaindb.com/docs/sql/quickstart) Graph Traverse aircraft rotations, connected legs and passenger itineraries.[Model relationships](https://originchaindb.com/docs/graph/quickstart) Vector Find related case notes using embeddings supplied by your application.[Search similar notes](https://originchaindb.com/docs/vector/quickstart) Full-text Match procedure terms, document titles and exact reference codes.[Find exact wording](https://originchaindb.com/docs/fts#walkthrough) Implementation path ## Start with one flight workflow. 1. 01 / Model ### Define IDs and relations. Register the source schemas and the connections your application needs to follow. 2. 02 / Index ### Prepare the search data. Write source rows, supply document embeddings and index selected text through separate calls. 3. 03 / Assemble ### Return useful evidence. Combine query results by ID, load source details and show their update time or document revision. Keep ingestion progress explicit so retries and source updates reach the relevant search entries. [See the ingestion pattern](https://originchaindb.com/docs/examples/atomic-multi-shape/kb-article) ## Give the team the connected records. [Build a first example](https://originchaindb.com/docs/quickstart)[Explore other industries](https://originchaindb.com/industries) --- # Multimodal data for banking | OriginChainDB Canonical source: https://originchaindb.com/industries/banking Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Banking · Customer and transaction context # Every case has a history. Keep the records connected. Bring account data, transaction relationships and policy search into one database for the applications your banking teams build. [Discuss your workload](https://originchaindb.com/contact)[Explore the quickstart](https://originchaindb.com/docs/quickstart) ## Build the case from its sources. Start with an account, inspect its transfers, and follow the recorded recipient. Your application brings that evidence together for an analyst. Account → transfers → counterparty → reviewIllustrative records Case context · account AC-104Open the transfer records ↓ Starting account AC-104Customer CU-28Account reference Two recorded transfers TX-21009:12 · outgoing From AC-104 To CP-08 Source payments / TX-210 TX-21109:18 · outgoing From AC-104 To CP-08 Source payments / TX-211 Shared recipient CP-08Harbor Supplies Two transfers resolve to the same recorded counterparty. Supporting policy evidenceRecipient review · version 3 Passage: “Check the recipient’s recorded details.” Policy PL-07 · passage 12 Application + analystReview before deciding. Compare the transfer sources with the relevant policy. Record the finding and the next action in the case. Illustrative account, transfers and policy text. The application retrieves the policy and assembles the case; the relationships do not establish a risk finding. ## Three places the context matters. 01 ### Customer servicing Bring an account's recent activity and related customer records into the service application, with the source IDs available for follow-up. SQL 02 ### Transaction investigations Move from a transfer to its recorded beneficiary and related accounts. The analyst interprets the connection and records the decision. SQL + graph 03 ### Internal policy retrieval Retrieve a policy by its exact wording or meaning. Show the passage, document version and source alongside the application's answer. Full-text + vector ## Use the query that fits the question. Share stable record IDs across your models so the application can resolve each result back to its source. SQL ### What happened on this account? Filter transactions by account and time. Join customer records and aggregate the result. [Query the records](https://originchaindb.com/docs/sql/quickstart) Graph ### Which records are connected? Declare account and counterparty references, then traverse the relationships with a bounded query. [Follow a relationship](https://originchaindb.com/docs/graph/quickstart) Full-text ### Where is this policy mentioned? Find exact terms or rank matching passages. Keep the document ID with every search result. [Search policy text](https://originchaindb.com/docs/fts) Vector ### Which passages discuss this issue? Retrieve similar text using your embeddings, then fetch the source passage for review. [Model the embeddings](https://originchaindb.com/docs/schemas/vector) Text indexes and vectors use their own write endpoints. Your ingestion pipeline manages their updates and records the source version. A focused first build ## Start with one investigation path. Use representative sample records before introducing customer data. [Review the design with us](https://originchaindb.com/contact) 1. 01 ### Define the records. Model accounts, transfers, counterparties and source documents with stable IDs. [Schema reference](https://originchaindb.com/docs/schemas) 2. 02 ### Test one complete case. Query the transfers, traverse their references, and check every returned source. Set result and traversal limits. [Graph quickstart](https://originchaindb.com/docs/graph/quickstart) 3. 03 ### Connect the review application. Keep credentials server-side. Apply the application's access rules and retain the analyst's decision with its evidence. [Authentication reference](https://originchaindb.com/docs/auth) --- # Multimodal data for contact centres | OriginChainDB Canonical source: https://originchaindb.com/industries/bpos-contact-centres Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) BPOs & contact centres # Give every agent the right context. Bring customer records, relevant cases, and searchable knowledge into the support applications your agents use. [Build the retrieval path](https://originchaindb.com/docs/rag)[Discuss your workflow](https://originchaindb.com/contact) From a question to its sources ## Put the evidence beside the answer. Your application retrieves authorized context, assembles the sources, and presents a draft for a human agent to review. Agent-assistance workflowIllustrative records, policy, and draft Case C-104 “My parcel arrived damaged. Can I get a replacement?” 01 / Customer contextSQL ### The current case Order O-52 Delivery Received Issue Damaged item Access Assigned agent Read only the records this agent is allowed to see. 02 / Retrieve knowledge ### Related, with a source Vector search`C-84` #### Similar resolved case Damaged parcel → photo reviewed → replacement approved. Full-text search`K-12` #### Exact policy passage “Request a photo of the damaged item before reviewing a replacement.” 03 / Application draft ### Ready for review “Please share a photo of the damaged item so we can review a replacement.” Policy source`K-12` The agent checks the policy and customer details before sending. Awaiting human review Application boundary Authorize access → retrieve records and passages → assemble context → generate a draft with your model → human review. OriginChainDB supplies the records and retrieval queries. Your application owns model calls, source references, workflow decisions, and the response sent to the customer. [Explore grounded retrieval](https://originchaindb.com/solutions/rag) Across the service desk ## Support the work between replies. 01 ### Carry context through a handoff. Store case notes, ownership, and recorded actions as structured rows. Fetch the relevant history when a new agent takes over, with programme and customer access checked by your application. The next agent seesWhat happened. What remains. 02 ### Find the policy behind the phrase. Use vector search for differently worded questions and full-text retrieval for exact terms. Combine candidates in your application and retain document IDs, versions, and passage locations. The answer includesA source the agent can inspect. 03 ### Give quality reviewers the full trail. Record the sources shown, the draft, agent edits, and the final disposition. Query those records for review queues and recurring issues; your team defines the assessment criteria. The reviewer comparesEvidence, response, and outcome. Build one assistance path first ## Start with the knowledge your agents trust. Use a reviewed set of support questions to check retrieval, permissions, and source freshness before broadening the workflow. [Follow the RAG guide](https://originchaindb.com/docs/rag) ### Model cases and source passages. Keep stable IDs for cases, documents, and chunks. Store policy versions alongside the text so your application can show which source it used. [Define the records](https://originchaindb.com/docs/schemas) ### Keep retrieval in step with changes. Write rows, vectors, and full-text entries through separate calls. Re-embed and re-index revised policies; test retries and removal of retired passages. [Set up semantic retrieval](https://originchaindb.com/docs/vector/quickstart) ### Test boundaries and missing answers. Verify programme isolation and agent permissions. When the available sources do not support a reply, have your application request clarification or escalate. [Review access controls](https://originchaindb.com/docs/auth) Built around your service workflow ## Better context starts with connected data. [Start with retrieval](https://originchaindb.com/docs/rag)[Talk through your use case](https://originchaindb.com/contact) --- # Multimodal data for capital markets | OriginChainDB Canonical source: https://originchaindb.com/industries/capital-markets Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Capital markets · Research and record context # Follow the instrument. Read the source. Connect instruments, issuers, events and research documents so analysts can inspect the records behind a question. [Discuss your research workflow](https://originchaindb.com/contact)[Explore the quickstart](https://originchaindb.com/docs/quickstart) ## An issuer is the link. A source is the evidence. Start with an instrument's recorded issuer. Inspect the linked event, filing and research note without losing the document version each passage came from. Instrument, issuer and research lineageIllustrative records One issuer. Several source records.Research context, with revisions InstrumentINS-104Instrument recordIssued by IssuerISS-18Example issuer 1. Issuer eventEVT-08Publication reference → DOC-21 2. Filing revisionDOC-21 · v2Retains the reference to v1 3. Research noteNOTE-06Cites DOC-21 · v1 The filing changed. The citation has a version.Inspect the illustrative revision trail DOC-21 / v1Original passage “Reporting period: Q1.” DOC-21 / v2Corrected passage “Reporting period: Q2.” NOTE-06Cited source: v1 The application can show the cited version alongside the correction for review. All identifiers and passages are illustrative. Source versions are records your ingestion pipeline maintains; they are not an automatic historical snapshot of every data model. ## Context for the analyst's next question. 01 / Issuer research ### What did the source actually say? Find a filing or research passage by its wording or meaning. Return the original document ID and revision with the result. [Search source text](https://originchaindb.com/docs/fts) 02 / Exposure context ### Which records point to this issuer? Join stored position and instrument records, then follow declared issuer relationships. Your application defines the scope and the analyst interprets it. [Query related records](https://originchaindb.com/docs/sql/quickstart) 03 / Source revisions ### Which version informed the note? Keep corrected documents associated with earlier versions. Store the source reference used for a note so a reviewer can reconstruct that context. [Model source versions](https://originchaindb.com/docs/schemas) A role for each model ## Query facts. Trace references. Retrieve passages. Use source IDs to connect the results in your research application. SQL ### Records and filters Filter instruments, join positions and inspect event timestamps. [SQL reference](https://originchaindb.com/docs/sql) Graph ### Entity relationships Traverse declared instrument, issuer and related-entity references. [Graph schema](https://originchaindb.com/docs/schemas/graph) Full-text ### Terms and phrases Find exact wording or rank matching passages in indexed source text. [Full-text reference](https://originchaindb.com/docs/fts) Vector ### Similar passages Retrieve research text using supplied embeddings, then inspect its source. [Vector schema](https://originchaindb.com/docs/schemas/vector) ## Make provenance part of the data model. Begin with a sample issuer, its instruments and a document that has been revised. 1. 01 ### Separate the timestamps. Store publication time, ingestion time and source revision as explicit fields. The application chooses which version a query should use. [Schema reference](https://originchaindb.com/docs/schemas) 2. 02 ### Index the source passages. Attach document and revision IDs to text and embeddings. Write each index through its own endpoint and track source updates in ingestion. [Index document text](https://originchaindb.com/docs/fts) 3. 03 ### Test the research trail. Check whether a result resolves to its original passage after a correction. Apply your application's data entitlements and retain the cited version. [Authentication reference](https://originchaindb.com/docs/auth) The database supplies record context and retrieval. Market interpretation and any resulting action remain with your application and analysts. [Review your data flow](https://originchaindb.com/contact) --- # Multimodal data for e-commerce and retail | OriginChainDB Canonical source: https://originchaindb.com/industries/ecommerce-retail Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Ecommerce & retail # Help shoppers find what fits. Connect product records, semantic search, keywords, and catalog relationships in one database for your storefront applications. [Build a product catalog](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog)[Discuss your storefront](https://originchaindb.com/contact?topic=retail) Discovery that knows the catalog ## One product. Several ways to find it. A shopper describes a need. Your application retrieves candidates, checks their catalog records, and decides what to show. From a shopper’s words to a useful productIllustrative catalog · application-orchestrated retrieval Shopper’s search “A daypack for wet commutes” Catalog record`sku_17` ### Commuter daypack Price $89 Inventory Available Category Bags 01Semantic matchVector Find product descriptions near the query’s embedding. Candidate ID `sku_17` 02Keyword intentFull-text Match words in the indexed product description. Waterproofdaypackfor daily travel. 03Category relationshipsGraph Follow the relationships you declare in the catalog. Commuter gearBags`sku_17` Your storefront application Merge candidate IDs → read catalog rows with SQL → check stock and price → rank results. Eligible productsOrdered by your merchandising rules The application combines separate retrieval results. Ranking, stock checks, and merchandising rules stay in your code; index filtering depends on the vector mode you choose. [Explore recommendations](https://originchaindb.com/solutions/recommendations) Across the shopping journey ## A useful next result. A useful next answer. 01 / Merchandising ### Make the shortlist yours. Find similar products with vectors. Follow declared category or accessory relationships with graph queries. Read price and availability from product rows before your application ranks the shortlist. - Exclude unavailable variants. - Apply regional catalog rules. - Choose how promotions affect ordering. [Build a tailored experience](https://originchaindb.com/solutions/personalization) 02 / Customer support ### Put the right policy beside the order. Retrieve product guides and return policies using keywords or semantic similarity. Fetch authorized order records separately, then give your support application the relevant context. Example question “Can I return the daypack I ordered?” Order statusReturn policyProduct guide [Connect retrieval to your assistant](https://originchaindb.com/docs/rag) Start with the product ID ## A catalog your team can maintain. Keep the ingestion path explicit so a new price, description, or product variant reaches the right query model. [Follow the catalog example](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog) 1. 01 ### Give each product a stable identity. Use the row’s primary key as the vector ID and full-text document ID. Declare category and supplier relations in the schema. [Define your schema](https://originchaindb.com/docs/schemas) 2. 02 ### Write each representation deliberately. Row, vector, and full-text updates are separate calls. Re-embed and re-index changed descriptions; retry failed steps. Declared graph edges follow the row write. [Index your embeddings](https://originchaindb.com/docs/vector/quickstart) 3. 03 ### Keep purchase rules close to checkout. Recheck inventory and price in your commerce system before purchase. Search relevance does not reserve stock or authorize access to a customer’s order. [Add keyword retrieval](https://originchaindb.com/docs/fts) Before the storefront rollout ## Evaluate with your own catalog. ### Test real search language. Compare category terms, product names, and descriptive queries against a shortlist your team has reviewed. ### Check change handling. Change a description, remove a product, and interrupt an indexing step. Verify the application handles missing or stale candidates. ### Measure the whole request. Include embedding generation, retrieval, row lookups, and application ranking when you measure response time. Build around your catalog ## Start with one discovery workflow. [Open the working example](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog)[Explore the database](https://originchaindb.com/product/database) --- # Multimodal data for payments | OriginChainDB Canonical source: https://originchaindb.com/industries/financial-services-payments Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Financial services & payments # Follow the payment. Keep the evidence in view. Connect processor events, merchant records and settlement evidence for the applications that investigate payment operations. [Discuss your payment workflow](https://originchaindb.com/contact)[Start with a SQL query](https://originchaindb.com/docs/sql/quickstart) ## Compare the sources. Keep the gap visible. A processor event and a settlement entry tell different parts of the story. Join their references and let the operations team inspect what is present or still missing. Payment and settlement evidenceIllustrative reconciliation worksheet Source AProcessor eventsPayment ID + provider reference Join by reference Source BSettlement entriesSettlement ID + provider reference REF-410Payment recordedSettlement recorded Payment recordPAY-42 Merchant M-18 Provider ref: REF-410 Settlement recordSET-88 Batch B-09 Provider ref: REF-410 What the join shows Both sources carry REF-410. The application can present the pair for comparison; matching a reference alone does not complete reconciliation. REF-411Payment recordedNo settlement in sample Payment recordPAY-43 Merchant M-18 Provider ref: REF-411 Settlement recordNo matching row Check the source feed and the selected batch. A question for the operator Has the settlement arrived in a later batch, or is source data missing? Keep the payment reference with the follow-up. Extend the investigation M-18Merchant relationships Source documentsReceipts, notes and policy versions Illustrative records. Queries expose source relationships; your application sets the reconciliation rules and the operator reviews exceptions. Operations in context ## From one exception to the records behind it. 01 / Reconciliation ### Inspect an unresolved reference. Compare payment and settlement rows by provider reference and batch. Your application routes unresolved records to an operator with both source IDs. [Query related rows](https://originchaindb.com/docs/sql/quickstart) 02 / Merchant context ### Follow the declared connections. Trace a payment to its merchant, terminal or account references. Inspect the recorded relationship before interpreting activity across those entities. [Traverse merchant relationships](https://originchaindb.com/docs/graph/quickstart) 03 / Case evidence ### Find the supporting document. Retrieve a receipt, support note or policy passage. Keep its original reference and version available to the person handling the case. [Search the text](https://originchaindb.com/docs/fts) Four models, specific jobs ## Match the query to the evidence. Use stable IDs to resolve every result back to the transaction, entity or document it represents. SQL Filter events, join references and aggregate a selected batch.[SQL reference](https://originchaindb.com/docs/sql) Graph Follow declared merchant, account and terminal relationships.[Relationship schema](https://originchaindb.com/docs/schemas/graph) Full-text Find exact references, phrases and ranked document matches.[Full-text reference](https://originchaindb.com/docs/fts) Vector Find similar issue descriptions using embeddings you supply.[Vector schema](https://originchaindb.com/docs/schemas/vector) ## Build a path an operator can inspect. Start with a representative payment batch and its settlement file. 1. 01 ### Retain source references. Store the provider reference, merchant ID, batch and event time with each record. Keep ingestion retries and ordering in your application's design. [Row API](https://originchaindb.com/docs/schemas/row-crud) 2. 02 ### Make updates explicit. Define the relationships you need. Write document text and vectors through their separate endpoints, with IDs that resolve to the source rows. [Schema reference](https://originchaindb.com/docs/schemas) 3. 03 ### Test the exception path. Include a missing settlement record and a corrected source. Check what the reviewer sees, which identity can access it, and how the decision is retained. [Authentication reference](https://originchaindb.com/docs/auth) OriginChainDB stores and queries the evidence. Your application owns reconciliation rules, payment actions and the review process. [Review your design](https://originchaindb.com/contact) --- # Multimodal data for healthcare | OriginChainDB Canonical source: https://originchaindb.com/industries/healthcare Sitemap last modified: 2026-09-23T08:33:23.000Z [← All industries](https://originchaindb.com/industries) Healthcare applications # Connect the records around each encounter. Build record exploration and knowledge workflows with encounter data, documents and care-team relationships. [Start with your data model](https://originchaindb.com/docs/schemas/tutorial)[Discuss your workflow](https://originchaindb.com/contact) ## A record with a way back to its context. Choose an encounter to see the source documents and team references connected to it. Encounter context explorerConceptual application data One patient recordConnected by a stable ID `patient:07` Selected encounter ### Visit record `encounter:12` Open the encounter fields and follow its document and team references. SQL · fieldsGraph · relationships Linked document Visit summary `document:31`Revision 2 Care-team reference Encounter team `team:08`Assigned in the source record Illustrative records. The application loads related data by ID and retains its source context. [Model the relationships](https://originchaindb.com/docs/graph/quickstart) Practical application workflows ## Make the next record easier to find. 01 / Explore ### Encounter workspaces Bring encounter fields, associated documents and team references into the same application view. [Follow record links](https://originchaindb.com/docs/graph/quickstart) 02 / Coordinate ### Referral operations Track a referral's status with its linked encounter, assigned team and follow-up task records. [Query workflow status](https://originchaindb.com/docs/sql/quickstart) 03 / Retrieve ### Service knowledge Search intake procedures and service manuals by meaning or exact terms, with source revisions attached. [Build knowledge retrieval](https://originchaindb.com/solutions/rag) Four complementary query models ## Fields. Relationships. Meaning. Words. Choose the query for the task, then load the original records by ID. [Explore the data models](https://originchaindb.com/product/database) [SQL Encounter fields and status Filter source, time and workflow attributes.](https://originchaindb.com/docs/sql/quickstart) [Graph Document and team links Traverse the relations your schema declares.](https://originchaindb.com/docs/graph/quickstart) [Vector Related reference material Search embeddings from your chosen model.](https://originchaindb.com/docs/vector/quickstart) [Full-text Exact document language Match titles, phrases and reference codes.](https://originchaindb.com/docs/fts#walkthrough) Implementation path ## Preserve the source as you connect the data. Start with one operational workflow and a clear model for record identity, updates and access. [Follow an ingestion example](https://originchaindb.com/docs/rag#ingest) 1. 01 ### Keep stable references. Register source IDs, encounter dates, document revisions and team relations in your schemas. 2. 02 ### Index explicitly. Write rows, supply embeddings and index selected text through separate calls; track updates and retries. 3. 03 ### Control what the application shows. Apply access rules before presenting records or retrieved text, and retain their source metadata. [Read the access guide](https://originchaindb.com/docs/auth) ## Start with an encounter. Build out the context. [Build a first example](https://originchaindb.com/docs/quickstart)[Explore other industries](https://originchaindb.com/industries) --- # Multimodal data for insurance | OriginChainDB Canonical source: https://originchaindb.com/industries/insurance Sitemap last modified: 2026-09-23T08:33:23.000Z [← All industries](https://originchaindb.com/industries) Insurance applications # Connect claims to their source evidence. Build investigation, policy knowledge and service workflows with connected records and searchable documents. [Start with the relationships](https://originchaindb.com/docs/graph/quickstart)[Discuss your workflow](https://originchaindb.com/contact) ## The record, the wording, the supporting file. Link a claim to the relevant policy version and supporting documents, retaining the source behind each result. A claim file with source contextIllustrative application data Claim file `claim:017` Start with the claim ### Keep the evidence attached. For reviewer context Structured record #### Claim 017 Policy reference `policy:042` Involved party `party:006` Service case `case:011` Policy sourceVersion 3 #### Policy wording `document:082` Stored source locationSection 4 Retain the original text and version with each retrieved passage. Supporting recordsLinked to `claim:017` Inspection note`document:018`Revision 1 Statement`document:026`Revision 2 Contact note`note:011`Revision 1 How the application assembles this file 1. SQL: load claim fields and policy references. 2. Graph: follow declared links to parties, documents and service cases. 3. Vector + full-text: retrieve related passages and exact terms separately, then load source records. [Explore the retrieval pattern](https://originchaindb.com/docs/rag) Illustrative records. Your application maintains the source IDs, policy versions and document links. From files to useful context ## Three workflows for insurance teams. Use the database to assemble information for your existing review and service processes. 01 ### Claims investigation Explore claims, involved parties and service providers alongside linked source documents. [Model the connected records](https://originchaindb.com/docs/graph/quickstart) 02 ### Policy knowledge Find wording by exact terms or related meaning, retaining the policy version and source section. [Build document retrieval](https://originchaindb.com/solutions/rag) 03 ### Service context Load policy and claim history with related communications for the servicing team. [Query the service records](https://originchaindb.com/docs/sql/quickstart) ## Choose how to find the evidence. [SQL ### Claim and policy fields Filter status, product, dates and source IDs.](https://originchaindb.com/docs/sql/quickstart) [Graph ### Parties and relationships Follow declared links between claims and involved entities.](https://originchaindb.com/docs/graph/quickstart) [Vector ### Related document passages Search embeddings supplied by your application.](https://originchaindb.com/docs/vector/quickstart) [Full-text ### Exact policy language Match phrases, document titles and reference codes.](https://originchaindb.com/docs/fts#walkthrough) Implementation path ## Make source references part of the model. 1. 01 / Register ### Define the claim context. Keep claim, policy, party and document IDs, plus policy versions and source locations. 2. 02 / Prepare ### Index selected documents. Write source rows, supply passage embeddings and index text through separate calls. 3. 03 / Retrieve ### Assemble the evidence. Combine result IDs in your application, apply access rules and show the original document references. Track source updates and index retries. Claim decisions remain part of your review process. [Follow the ingestion pattern](https://originchaindb.com/docs/rag#ingest) ## Start with a claim. Keep its sources in view. [Build a first example](https://originchaindb.com/docs/quickstart)[Explore other industries](https://originchaindb.com/industries) --- # Multimodal data for lending | OriginChainDB Canonical source: https://originchaindb.com/industries/lending Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Lending · Application evidence # An application is more than its form fields. Bring borrower records, business relationships and supporting documents into the applications your lending teams use to review a case. [Discuss your lending workflow](https://originchaindb.com/contact)[Build a sample application](https://originchaindb.com/docs/quickstart) ## Give every piece of evidence a source. Resolve a document passage or a declared business relationship back to the application it supports. The reviewer can then inspect the record behind the result. From application to review packetIllustrative business-loan application Application fileAPP-17 01 / Application fields Applicant CU-92 Business BIZ-18 Purpose Working capital Source applications / APP-17 02 / Declared relationship CU-92Applicant Director of BIZ-18Business Source: declaration REL-06. The recorded link is available for inspection. 03 / Document passage Business statementDOC-31 · version 2 · passage 4 “Prepared for BIZ-18. Statement period: April–June.” Illustrative extracted text. Resolve DOC-31 to the original source before using the passage. Application assembles ### A packet for review. 1. 01Field values and their source IDs 2. 02Declared applicant relationships 3. 03Document passages and versions Reviewer interprets the evidence. Eligibility rules and the lending decision belong to your application and review process. Illustrative records and document text. The database retrieves fields, relationships and passages; it does not verify the applicant’s declarations. Across the application lifecycle ## Keep context with the case. The same borrower and application IDs can connect intake, document review and later servicing. 01 ### Underwriting operations Compare submitted fields with the evidence received. Surface missing references, document versions and outstanding review tasks in your application. [Query application records](https://originchaindb.com/docs/sql/quickstart) 02 ### Document retrieval Search for a specific term or retrieve passages related to a reviewer's question. Show the original document and its version with the answer. [Search document passages](https://originchaindb.com/docs/fts) 03 ### Borrower servicing Connect later correspondence and case notes to the borrower's application history, so the servicing team can inspect the relevant context. [Define record relationships](https://originchaindb.com/docs/schemas/graph) ## Different evidence. A query for each. Resolve model results by their source IDs rather than treating a search match as a verified fact. SQL ### Structured facts Filter application fields, join borrower records and inspect document metadata. [SQL reference](https://originchaindb.com/docs/sql) Graph ### Declared relationships Follow recorded links between applicants, businesses and related application records. [Graph quickstart](https://originchaindb.com/docs/graph/quickstart) Full-text ### Specific wording Find exact terms and phrases in the text your ingestion pipeline indexes. [Full-text reference](https://originchaindb.com/docs/fts) Vector ### Related passages Use supplied embeddings to retrieve text by similarity, then read the source. [Vector schema](https://originchaindb.com/docs/schemas/vector) ## Start with one complete review packet. Use sample records to test the path from ingestion to reviewer hand-off. 1. 01 ### Model the references. Give applications, borrowers, businesses and documents stable IDs. Record the source and version of extracted fields. [Schema reference](https://originchaindb.com/docs/schemas) 2. 02 ### Keep indexing explicit. Write passages and embeddings through their separate endpoints. Have the ingestion pipeline track updated documents and stale results. [Text indexing](https://originchaindb.com/docs/fts) 3. 03 ### Evaluate the review path. Test a missing document, an updated version and a restricted identity. Keep credentials server-side and retain the reviewer's source references. [Authentication reference](https://originchaindb.com/docs/auth) Your application owns eligibility rules and lending decisions. OriginChainDB stores and queries the records that support that process. [Review the design](https://originchaindb.com/contact) --- # Multimodal data for mutual funds | OriginChainDB Canonical source: https://originchaindb.com/industries/mutual-funds Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Mutual funds # Connect the holding. Keep the source. Bring fund records, issuer relationships, and searchable disclosures into your research and investor-service applications. [Model connected records](https://originchaindb.com/docs/graph/quickstart)[Discuss your data workflow](https://originchaindb.com/contact) A path from position to publication ## Know what connects—and which version you read. Follow declared holding and issuer relationships, then retrieve the disclosure passages relevant to the analyst’s question. Fund research contextIllustrative records and application-selected versions Research question Which source documents relate to this fund’s holdings? 1. FundResearch fund`F-01`has holding 2. HoldingEquity position`H-21`issued by 3. IssuerNorthbridge Industries`I-07`has disclosure 4. DisclosureIssuer report`D-12`Source ID retained Document history ### Same source. A new revision. Your application decides which version is eligible for retrieval. Version 2Archived Preserved for earlier research records. `D-12:v2` Version 3Selected for review Passages linked to this document version. `D-12:v3` Application assembles - Holding and issuer records - Matching disclosure passages - Document ID, version, and location Relationships provide context; they do not establish investment merit. The application combines separate queries, selects source versions, and leaves interpretation to the analyst. [Build source-linked retrieval](https://originchaindb.com/docs/rag) Useful context across the organisation ## Research, portfolio context, and service. 01 / Research ### Find the passage behind the question. Use semantic retrieval for differently worded topics and full-text search for named entities or exact terms. Retain the document version and passage location when presenting the results. [Search disclosure text](https://originchaindb.com/docs/fts) 02 / Portfolio context ### Trace the records behind a holding. Query fund and position rows, then follow issuer relationships. Store effective dates in your model so the application can distinguish a historical portfolio from the latest imported records. [Follow relationships](https://originchaindb.com/docs/graph) 03 / Investor service ### Retrieve the relevant fund document. Find passages in factsheets, service instructions, and fund disclosures. Combine them with authorized service records in your application, keeping the source visible for a staff member to review. [Review access controls](https://originchaindb.com/docs/auth) Define the data before the answer ## Give each model a clear job. Start with one fund, a reviewed set of holdings, and versioned source documents. Record both source dates and ingestion times; test missing and superseded material before expanding coverage. SQL / records ### Keep identities and dates explicit. Model fund IDs, holdings, issuers, and source versions as rows. Query the fields your application uses to select an appropriate reporting period. [Define the schema](https://originchaindb.com/docs/schemas) Graph / relationships ### Declare the paths you need. Connect holding records to funds and issuers using schema relations. Traverse those paths to gather related records. [Create relationships](https://originchaindb.com/docs/graph/quickstart) Vector / meaning ### Index disclosure passages. Generate embeddings in your pipeline. Keep vector IDs tied to passage rows, and evaluate results against questions your analysts have reviewed. [Set up semantic search](https://originchaindb.com/docs/vector/quickstart) Full-text / terminology ### Keep words and sources together. Index text separately for keyword retrieval. When a document changes, refresh its rows, embeddings, and text entries through their separate calls. [Add full-text search](https://originchaindb.com/docs/fts) Start with a source you can verify ## Build a research path worth following. [Open the retrieval guide](https://originchaindb.com/docs/rag)[Discuss your use case](https://originchaindb.com/contact) --- # Multimodal data for NBFCs | OriginChainDB Canonical source: https://originchaindb.com/industries/nbfcs Sitemap last modified: 2026-09-23T08:33:23.000Z [← All industries](https://originchaindb.com/industries) Non-banking financial companies # Borrower records. Business context. Connect borrower, loan, business and branch data with the documents your operations team needs. [Start with a portfolio query](https://originchaindb.com/docs/sql/quickstart)[Discuss your workflow](https://originchaindb.com/contact) ## See the business behind the record. Start with a borrower ID, then follow business links, branch context and supporting documents. Borrower and portfolio contextIllustrative application records Start with a borrower ### Borrower 014 `borrower:014` Linked loan recordloan:028 Keep the source fields and follow declared relationships. Business link #### Business profile `business:008` Company fields and related-party references. Graph + SQL Servicing branch #### Branch context `branch:004` Portfolio grouping and assigned service records. SQL Supporting source #### Document reference `document:031` Original text, source location and document version. Vector + Full-text One starting record, several query paths. Your application combines the results by ID. Supplied embeddings and searchable text are indexed separately. Example: group borrower records by branch Illustrative SQL for a schema with a `branch_id` field. ``` SELECT branch_id, COUNT(*) AS borrowers FROM borrowers GROUP BY branch_id; ``` [Build a portfolio query](https://originchaindb.com/docs/sql/quickstart) Three operational questions ## Find the records. Understand the connections. 01 ### Who is connected? Explore declared borrower, business, guarantor and partner relationships with their source records. [Relationship review](https://originchaindb.com/docs/graph/quickstart) 02 ### What is in this portfolio? Filter and group loan records by branch, product or recorded status, then inspect individual rows. [Portfolio queries](https://originchaindb.com/docs/sql/quickstart) 03 ### Where is the guidance? Find relevant servicing procedures and prior case notes with their original source references. [Servicing knowledge](https://originchaindb.com/solutions/rag) Four query models ## One application. Several ways to explore. OriginChainDB stores and queries the representations. Your application chooses the queries and assembles the results. [Explore the database](https://originchaindb.com/product/database) SQL Structured borrower and loan fields, branch filters and portfolio aggregates.[Query records](https://originchaindb.com/docs/sql/quickstart) Graph Declared relationships between borrowers, businesses, branches and partners.[Follow relationships](https://originchaindb.com/docs/graph/quickstart) Vector Related servicing notes and document passages, using supplied embeddings.[Search by similarity](https://originchaindb.com/docs/vector/quickstart) Full-text Exact phrases, document titles and reference codes in indexed text.[Match the wording](https://originchaindb.com/docs/fts#walkthrough) Implementation path ## Start with one branch or workflow. 1. 01 / Connect the sources ### Keep stable identifiers. Register borrower, loan, business and document schemas; retain source IDs and declared relations. 2. 02 / Prepare retrieval ### Track the separate writes. Save rows, supply embeddings and index selected document text through their own endpoints. 3. 03 / Build the view ### Return records with context. Apply access rules, combine result IDs and show source versions and update times. Keep ingestion progress and retries explicit when a source record or document changes. [See the ingestion pattern](https://originchaindb.com/docs/examples/atomic-multi-shape/kb-article) ## Build the view your operations team needs. [Build a first example](https://originchaindb.com/docs/quickstart)[Explore lending workflows](https://originchaindb.com/industries/lending) --- # Multimodal data for telecommunications | OriginChainDB Canonical source: https://originchaindb.com/industries/telecom Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Telecommunications # See the service behind the circuit. Connect network inventory, service relationships, and operational knowledge for investigation and customer-support applications. [Model the dependencies](https://originchaindb.com/docs/graph/quickstart)[Discuss your network data](https://originchaindb.com/contact) Follow the recorded relationships ## From one circuit to its service context. Model the dependencies your team knows, then retrieve the records and runbooks an investigator needs. Service dependency mapIllustrative inventory · not live network telemetry Investigation question Which services depend on circuit C-14? SiteNorth exchange`S-03` CircuitAccess circuit`C-14` ServiceBranch connectivity`SV-09` Customer accountHarbor retail`A-27` Declared dependencies: site → circuit → service → account Incident record`T-219` ### Intermittent connectivity Reported against C-14. Keep the recorded symptoms, timestamps, and investigation notes together. SQL · structured context Runbook source`R-07` ### Access-circuit checks Retrieve passages by exact terms or similar descriptions of the reported problem. Full-text + vector candidates Your investigation application ### A source-linked case view - Related service and account - Ticket and source records - Runbook passages for review Your application combines these separate queries. A stored dependency identifies context for review; it does not establish that a circuit caused an incident or that a customer is currently affected. [Explore graph queries](https://originchaindb.com/docs/graph) For operations and service teams ## Keep the investigation connected. Network investigation ### Trace the relevant dependencies. Start from a site, circuit, or service ID. Traverse declared relationships, fetch the linked records, and review related tickets. Your ingestion process supplies inventory updates and observations from existing systems. FindRelated assets and services ReviewEvidence, ownership, and freshness [Query connected records](https://originchaindb.com/docs/examples/graph/neighbors) Service support ### Bring the runbook into the case. Look up authorized account and service details with SQL. Use keyword and semantic retrieval to find troubleshooting guidance. Show source versions so an engineer can assess the next step before taking action. RetrieveCurrent guidance and prior cases PresentContext an agent can verify [Build knowledge retrieval](https://originchaindb.com/docs/rag) Four models, defined responsibilities ## Build from your inventory outward. Choose stable asset IDs and record when each dependency was observed. Keep retrieval orchestration and operational actions in your application. Write rows, vectors, and full-text entries separately. Re-index changed runbooks and verify access before passing customer context to a model or agent. SQL ### Read structured facts. Model inventory, cases, service ownership, and source versions. [Define the schema](https://originchaindb.com/docs/schemas) Graph ### Follow known dependencies. Declare relationships between your records and choose the traversal depth for each investigation. [Create a relationship](https://originchaindb.com/docs/graph/quickstart) Vector ### Find similarly described issues. Generate embeddings in your pipeline; test retrieval against cases your team has reviewed. [Index embeddings](https://originchaindb.com/docs/vector/quickstart) Full-text ### Keep exact terms searchable. Index runbook passages for product names, fault descriptions, and operational terminology. [Add keyword retrieval](https://originchaindb.com/docs/fts) Start with a known service path ## Make one investigation easier to follow. [Open the graph quickstart](https://originchaindb.com/docs/graph/quickstart)[Discuss your use case](https://originchaindb.com/contact) --- # Multimodal data for wealth management | OriginChainDB Canonical source: https://originchaindb.com/industries/wealth-management Sitemap last modified: 2026-09-23T08:33:23.000Z [All industries](https://originchaindb.com/industries) Wealth management · Adviser context # Client context. With the sources attached. Connect client records, holdings, household relationships and research documents for the applications your advisers use to prepare and follow up. [Discuss your adviser workflow](https://originchaindb.com/contact)[Explore the quickstart](https://originchaindb.com/docs/quickstart) ## Prepare the conversation from its records. Scope the client, retrieve the relevant holdings and documents, and keep each source reference with the briefing your application assembles. Client scope → source records → adviser briefingIllustrative preparation flow Prepare a briefingClient CU-07Source references stay attached 01 / Client scope CU-07 Household reference HH-04 Account reference AC-12 The application selects the client and the records it is allowed to retrieve. 02 / Retrieve the working context ### Mandate M-09 · version 3 Recorded client instructions and their source document. Linked client: CU-07 ### Holdings Import H-14 Stored account and instrument rows, with the supplied as-of date. Linked account: AC-12 ### Research R-04 · version 2 Retrieved passages with document and instrument references. Source passage: R-04 / 3 03 / Application assemblesBrief BR-05 Client context, source versions and questions for the adviser. What should the adviser inspect? Mandate sourceCheck the instruction document and the version referenced. Holdings contextCheck the imported dataset and its supplied as-of date. Research evidenceRead the source passage before using it in the conversation. Illustrative records and workflow. The application assembles the briefing; source retrieval does not establish suitability or produce an investment recommendation. ## Keep the working context close. 01 / Servicing ### Pick up the client's history. Bring recorded preferences, correspondence and mandate versions into the servicing application. Link follow-up notes to the client and conversation. [Query client records](https://originchaindb.com/docs/sql/quickstart) 02 / Research ### Find the relevant source passage. Search research text by term or meaning. Return the document and revision so the adviser can inspect where a passage came from. [Explore document search](https://originchaindb.com/docs/fts) 03 / Preparation ### Assemble an inspectable briefing. Resolve household, account and instrument references. Your application combines the retrieved context and leaves interpretation with the adviser. [Follow record relationships](https://originchaindb.com/docs/graph/quickstart) Choose the query by the question ## Facts, relationships and source passages. SQL Which records belong in this briefing? Filter client and account rows, join stored holdings and inspect their source timestamps. [SQL reference](https://originchaindb.com/docs/sql) Graph How are the entities connected? Traverse declared household, account, instrument and issuer relationships. [Graph schema](https://originchaindb.com/docs/schemas/graph) Full-text Where does the document say that? Find exact phrases or rank matching passages in text you have indexed. [Full-text reference](https://originchaindb.com/docs/fts) Vector Which passages discuss this topic? Retrieve text using supplied embeddings and return its original source ID. [Vector schema](https://originchaindb.com/docs/schemas/vector) A focused first build ## Start with one client briefing. Use sample client records, a holdings import and a versioned mandate to test the full retrieval path. [Review the design with us](https://originchaindb.com/contact) 1. 01 ### Retain identity and provenance. Store client, household and account IDs alongside source IDs, document versions and the supplied holdings date. [Schema reference](https://originchaindb.com/docs/schemas) 2. 02 ### Keep retrieval sources aligned. Index text and embeddings through their separate endpoints. Have your ingestion pipeline track revisions and resolve results back to source records. [Text indexing](https://originchaindb.com/docs/fts) 3. 03 ### Review the client boundary. Apply your application's access rules across every query path. Test restricted users, outdated documents and missing references before introducing client data. [Authentication reference](https://originchaindb.com/docs/auth) The database retrieves context. Your application and advisers own interpretation, recommendations and client decisions. --- # Newsroom: announcements & ideas | OriginChainDB Canonical source: https://originchaindb.com/newsroom Sitemap last modified: 2026-09-28T07:29:43.000Z From the team at OriginChainDB # Inside OriginChainDB What we’re building. Who we’re building with. Where data goes next. [Follow the newsroom](https://originchaindb.com/newsroom/rss.xml) Product updates28 Sept 2026 ## [More room to build. From $29 a month.](https://originchaindb.com/newsroom/pooled-configurations) Meet our 25 GB and 50 GB pooled configurations for building with OriginChainDB. [Read the announcement](https://originchaindb.com/newsroom/pooled-configurations) 01 / From OriginChainDB ## The latest from us. [Release notes](https://originchaindb.com/changelog) Product updates28 Sept 2026 ### [More room to build. From $29 a month.](https://originchaindb.com/newsroom/pooled-configurations) Meet our 25 GB and 50 GB pooled configurations for building with OriginChainDB. Go deeper [Product releasesFeatures and fixes, version by version.](https://originchaindb.com/changelog)[Engineering journalHow the database works.](https://originchaindb.com/blogs)[Meet OriginChainDBOur company and the idea behind it.](https://originchaindb.com/company) 02 / What’s next ## The next chapter starts here. New product announcements, company milestones and collaborations. We’ll share the details here when they’re ready. [Follow future announcements](https://originchaindb.com/newsroom/rss.xml) 03 / Collaborations ## Good work starts together. Build with OriginChainDB ### A shared question. Something worth building. Have an integration, research project or industry use case in mind? Let’s explore it together. [Discuss a collaboration](https://originchaindb.com/contact?topic=partner) Confirmed collaborations will be announced here. 04 / Our perspective ## The industry. Through our lens. What it means for the systems you build. AI infrastructure27 Sept 2026 ### [MCP’s roadmap puts agent identity in focus](https://originchaindb.com/newsroom/mcp-roadmap-agent-data-access) A practical reading of the updated MCP roadmap for teams connecting agents to business data. [Read our perspective](https://originchaindb.com/newsroom/mcp-roadmap-agent-data-access) Database releases27 Sept 2026 ### [PostgreSQL 19 Beta 4: a moment to test your workload](https://originchaindb.com/newsroom/postgresql-19-beta-evaluation) Treat a database beta as an opportunity to validate application behavior before planning a production change. [Read our perspective](https://originchaindb.com/newsroom/postgresql-19-beta-evaluation) Industry reading listHeadlines from official sources · checked 30 Sept 2026 PostgreSQL · 29 Sept 2026 ### [Dasha 1.8: index recommendations, I/O analysis, schema checks and log insights](https://www.postgresql.org/about/news/dasha-18-index-recommendations-io-analysis-schema-checks-and-log-insights-3387/) PostgreSQL · 28 Sept 2026 ### [PL/Haskell v6.0 Released](https://www.postgresql.org/about/news/plhaskell-v60-released-3388/) PostgreSQL · 28 Sept 2026 ### [pgEdge Announces pgEdge Starfleet, a New Postgres Cloud Platform to Bridge the AI Prototype to Production Chasm](https://www.postgresql.org/about/news/pgedge-announces-pgedge-starfleet-a-new-postgres-cloud-platform-to-bridge-the-ai-prototype-to-production-chasm-3389/) PostgreSQL · 24 Sept 2026 ### [PostgreSQL 19 Beta 4 Released!](https://www.postgresql.org/about/news/postgresql-19-beta-4-released-3386/) Model Context Protocol · 22 Aug 2026 ### [The New MCP Roadmap](https://blog.modelcontextprotocol.io/posts/mcp-roadmap/) Model Context Protocol · 28 Jul 2026 ### [The 2026-07-28 Specification](https://blog.modelcontextprotocol.io/posts/2026-07-28/) Model Context Protocol · 27 Jul 2026 ### [The Official Ruby SDK for MCP Reaches 1.0](https://blog.modelcontextprotocol.io/posts/ruby-sdk-1-0/) Model Context Protocol · 29 Jun 2026 ### [Beta SDKs for the 2026-07-28 MCP Spec Release Candidate Are Here](https://blog.modelcontextprotocol.io/posts/sdk-betas-2026-07-28/) Automatically refreshed source headlines. These are external publications, not OriginChainDB announcements or partnerships. A story worth sharing? ## Let’s talk. Company news, press enquiries or a project we could build together. [Contact the team](https://originchaindb.com/contact?topic=other)[Newsroom RSS](https://originchaindb.com/newsroom/rss.xml) --- # MCP’s roadmap puts agent identity in focus | OriginChainDB Newsroom Canonical source: https://originchaindb.com/newsroom/mcp-roadmap-agent-data-access Sitemap last modified: 2026-09-28T10:36:59.000Z Industry perspectives / AI infrastructure # MCP’s roadmap puts agent identity in focus A practical reading of the updated MCP roadmap for teams connecting agents to business data. [Read our perspective](https://originchaindb.com/newsroom/mcp-roadmap-agent-data-access#story)[Original announcement](https://blog.modelcontextprotocol.io/posts/mcp-roadmap/) [OriginChainDB Team](https://originchaindb.com/company)27 September 2026 [Image: An agent identity passes through a scoped permission boundary before accessing data.] [Image: An agent identity passes through a scoped permission boundary before accessing data.] An OriginChainDB perspective. 01 ## Identify the caller Separate the agent from the person it represents. 02 ## Limit its authority Match permissions to the task and dataset. 03 ## Inspect the outcome Record the action, result and authorization context. [← All newsroom stories](https://originchaindb.com/newsroom) ## What changed The MCP maintainers’ August roadmap identifies agent identity, transport hardening and improved result contracts among its priorities. It distinguishes work already released from proposals for future versions. [Read the maintainers’ roadmap](https://blog.modelcontextprotocol.io/posts/mcp-roadmap/). ## Our reading Connecting an agent to a database is an authorization decision as well as an integration task. A tool’s presence does not establish which records its caller should be allowed to read or change. For a database integration, document three identities: the person granting access, the agent making the request, and the credential accepted by the data service. Check where their permissions are enforced rather than assuming the tool protocol supplies that boundary. ## Before connecting production data - Start with the smallest useful read scope. - Test an explicitly forbidden operation as well as a permitted one. - Check credential expiry and revocation behavior. These are evaluation questions, not claims that every roadmap feature is implemented by OriginChainDB. Use our [integration documentation](https://originchaindb.com/docs/sdk) to confirm the interfaces available for your application. Primary source · Published 22 August 2026 [Model Context Protocol: The New MCP Roadmap ↗](https://blog.modelcontextprotocol.io/posts/mcp-roadmap/) Our interpretation is separate from the publisher’s announcement. External product claims are not OriginChainDB benchmark results. [Suggest a correction →](https://originchaindb.com/contact?topic=other) From reading to building ## Put your data to work. Explore SQL, vector, graph and full-text in one database. [Open the quickstart](https://originchaindb.com/docs/quickstart) ## More from OriginChainDB [Visit the newsroom ↗](https://originchaindb.com/newsroom) [[Image: OriginChainDB pooled configurations: 25 GB for $29 per month with 20 GB transfer, and 50 GB for $49 per month with 50 GB transfer.] [Image: OriginChainDB pooled configurations: 25 GB for $29 per month with 20 GB transfer, and 50 GB for $49 per month with 50 GB transfer.] Product updates ### More room to build. From $29 a month. Read story ↗](https://originchaindb.com/newsroom/pooled-configurations)[[Image: A baseline and a candidate database are evaluated with the same representative queries to compare behavior.] [Image: A baseline and a candidate database are evaluated with the same representative queries to compare behavior.] Industry perspectives ### PostgreSQL 19 Beta 4: a moment to test your workload Read story ↗](https://originchaindb.com/newsroom/postgresql-19-beta-evaluation) --- # OriginChainDB pooled database configurations: 25 GB and 50 GB | OriginChainDB Newsroom Canonical source: https://originchaindb.com/newsroom/pooled-configurations Sitemap last modified: 2026-09-28T10:36:59.000Z Product updates / Product # More room to build. From $29 a month. Meet our 25 GB and 50 GB pooled configurations for building with OriginChainDB. [Compare configurations](https://originchaindb.com/pricing)[Read the announcement](https://originchaindb.com/newsroom/pooled-configurations#story) [OriginChainDB Team](https://originchaindb.com/company)28 September 2026 [Image: OriginChainDB pooled configurations: 25 GB for $29 per month with 20 GB transfer, and 50 GB for $49 per month with 50 GB transfer.] [Image: OriginChainDB pooled configurations: 25 GB for $29 per month with 20 GB transfer, and 50 GB for $49 per month with 50 GB transfer.] Built for your next application. 01 ## Choose your storage 25 GB at $29/month or 50 GB at $49/month. 02 ## Build on shared compute Single-node configurations in Mumbai. 03 ## Connect your application Use the console to choose a configuration and get started. [← All newsroom stories](https://originchaindb.com/newsroom) ## Two configurations. One place to start. Our pooled configurations give teams two storage options: 25 GB for $29 per month and 50 GB for $49 per month. Both use shared compute on a single node in Mumbai, with monthly billing in USD. The 25 GB configuration includes 20 GB of monthly data transfer. The 50 GB configuration includes 50 GB. See [current pricing](https://originchaindb.com/pricing) for the configuration details before ordering. ## Know what you’re choosing Pooled configurations do not include high availability, replicas or point-in-time recovery. Dedicated configurations serve workloads that need those options. ## Start with your workload Explore the [database](https://originchaindb.com/product/database), follow the [quickstart](https://originchaindb.com/docs/quickstart), or [open the console](https://app.originchaindb.com/instances) to choose a configuration. Product reference [OriginChainDB configuration pricing ↗](https://originchaindb.com/pricing) [Suggest a correction →](https://originchaindb.com/contact?topic=other) Make room for your next idea ## Choose your starting point. Compare configurations, then get started in the console. [Explore pricing](https://originchaindb.com/pricing) ## More from OriginChainDB [Visit the newsroom ↗](https://originchaindb.com/newsroom) [[Image: An agent identity passes through a scoped permission boundary before accessing data.] [Image: An agent identity passes through a scoped permission boundary before accessing data.] Industry perspectives ### MCP’s roadmap puts agent identity in focus Read story ↗](https://originchaindb.com/newsroom/mcp-roadmap-agent-data-access)[[Image: A baseline and a candidate database are evaluated with the same representative queries to compare behavior.] [Image: A baseline and a candidate database are evaluated with the same representative queries to compare behavior.] Industry perspectives ### PostgreSQL 19 Beta 4: a moment to test your workload Read story ↗](https://originchaindb.com/newsroom/postgresql-19-beta-evaluation) --- # PostgreSQL 19 Beta 4: a moment to test your workload | OriginChainDB Newsroom Canonical source: https://originchaindb.com/newsroom/postgresql-19-beta-evaluation Sitemap last modified: 2026-09-28T10:36:59.000Z Industry perspectives / Database releases # PostgreSQL 19 Beta 4: a moment to test your workload Treat a database beta as an opportunity to validate application behavior before planning a production change. [Read our perspective](https://originchaindb.com/newsroom/postgresql-19-beta-evaluation#story)[Original announcement](https://www.postgresql.org/about/news/postgresql-19-beta-4-released-3386/) [OriginChainDB Team](https://originchaindb.com/company)27 September 2026 [Image: A baseline and a candidate database are evaluated with the same representative queries to compare behavior.] [Image: A baseline and a candidate database are evaluated with the same representative queries to compare behavior.] An OriginChainDB perspective. 01 ## Establish a baseline Save representative queries and expected results. 02 ## Test in isolation Use a separate environment and representative data. 03 ## Review the differences Compare behavior, plans and application latency. [← All newsroom stories](https://originchaindb.com/newsroom) ## What changed The PostgreSQL Global Development Group announced PostgreSQL 19 Beta 4 on September 24. It is a beta release; its announcement should not be read as a production-upgrade recommendation. [Read the official release notice](https://www.postgresql.org/about/news/postgresql-19-beta-4-released-3386/). ## Our reading A release milestone is a useful prompt to improve your own test corpus. Collect the queries your application actually runs, including filters, joins, pagination and error cases. Record expected results before looking at performance. Then compare application-level latency using equivalent datasets and connection behavior. A fast query measured inside a database is not the same measurement as a request that includes networking and serialization. ## Make the results reusable Save the database version, schema, query parameters and test conditions with each result. That record is useful for future upgrades and for evaluating a different database. OriginChainDB is a separate database engine. Client-protocol compatibility does not imply complete PostgreSQL feature compatibility. Start with the [SQL documentation](https://originchaindb.com/docs/sql/quickstart) when evaluating a workload. Primary source · Published 24 September 2026 [PostgreSQL Global Development Group: PostgreSQL 19 Beta 4 ↗](https://www.postgresql.org/about/news/postgresql-19-beta-4-released-3386/) Our interpretation is separate from the publisher’s announcement. External product claims are not OriginChainDB benchmark results. [Suggest a correction →](https://originchaindb.com/contact?topic=other) From reading to building ## Put your data to work. Explore SQL, vector, graph and full-text in one database. [Open the quickstart](https://originchaindb.com/docs/quickstart) ## More from OriginChainDB [Visit the newsroom ↗](https://originchaindb.com/newsroom) [[Image: OriginChainDB pooled configurations: 25 GB for $29 per month with 20 GB transfer, and 50 GB for $49 per month with 50 GB transfer.] [Image: OriginChainDB pooled configurations: 25 GB for $29 per month with 20 GB transfer, and 50 GB for $49 per month with 50 GB transfer.] Product updates ### More room to build. From $29 a month. Read story ↗](https://originchaindb.com/newsroom/pooled-configurations)[[Image: An agent identity passes through a scoped permission boundary before accessing data.] [Image: An agent identity passes through a scoped permission boundary before accessing data.] Industry perspectives ### MCP’s roadmap puts agent identity in focus Read story ↗](https://originchaindb.com/newsroom/mcp-roadmap-agent-data-access) --- # Partner with OriginChainDB | OriginChainDB Canonical source: https://originchaindb.com/partners Sitemap last modified: 2026-09-24T06:54:16.000Z partners # Partner with OriginChainDB Build integrations, deliver database deployments or bring OriginChainDB to customers through a technology or commercial partnership. [Talk about partnering](https://originchaindb.com/contact?topic=partner) [Read the docs](https://originchaindb.com/docs) 01 · memberships ## Programmes we are part of. membership ### NVIDIA Inception Program member OriginChainDB is a member of the NVIDIA Inception program. We're working with NVIDIA on the GPU-accelerated paths in our roadmap - cuVS for very large vector indexes, Triton Inference Server for managed embedding and reranker hosting, and TensorRT-LLM for a future managed LLM offering. [NVIDIA Inception →](https://www.nvidia.com/en-us/startups/) membership ### AI Council of India (IAMAI) member Member of the AI Council of India, convened by the Internet and Mobile Association of India (IAMAI). [IAMAI →](https://www.iamai.in/) © 2026 NVIDIA, the NVIDIA logo, and NVIDIA Inception are trademarks and/or registered trademarks of NVIDIA Corporation in the U.S. and other countries. 02 · partner programme ## Three ways to work with us. ### Systems integrators & consultancies You design and deliver AI systems for enterprises; we supply the database layer, engineering support during delivery and a shared channel for the life of the account. ### Technology partners Embedding and reranking model providers, agent frameworks, observability and data-platform vendors. Joint reference architectures and documented integrations. ### Resellers & marketplaces Procurement through an existing vendor relationship or a cloud marketplace listing, where a customer's purchasing process requires it. ## Delivering an AI system for a client? Start with the database. Tell us about the engagement. We bring an engineer to the scoping call. [Talk about partnering](https://originchaindb.com/contact?topic=partner) [Start free](https://app.originchaindb.com/signup) --- # Evaluate your workload in a pilot | OriginChainDB Canonical source: https://originchaindb.com/pilot Sitemap last modified: 2026-09-23T08:33:23.000Z technical pilot # Evaluate your workload in a pilot Agree on success criteria, test your workload on a dedicated instance and review the results with our team. [Request a pilot](https://originchaindb.com/contact?topic=pilot) [Book a technical walkthrough first](https://originchaindb.com/contact?topic=walkthrough) 01 · the four weeks ## Scope, load, test, report. 1. Week 1 ### Scope Data shapes, query patterns, expected volumes, availability target, region. We write the success criteria together — latency, recall, throughput, cost — and agree how each will be measured. 2. Week 2 ### Load A dedicated instance is provisioned in your region (under two minutes). Your data lands through the SDK or HTTP API; schemas, indexes and the natural-language layer are configured against real queries. 3. Week 3 ### Integrate & test Your service talks to the instance. We run the agreed benchmarks side by side with your current stack, on the same data, and record the numbers. 4. Week 4 ### Report A written report: every criterion, the measured result, the configuration that produced it and its monthly price. You decide with numbers, not a demo. you bring - → A representative slice of the data (or a generator that produces its shape). - → Ten to twenty real queries, including the slow ones. - → The numbers that matter to you and the thresholds that make the pilot a pass. - → One engineer who can wire a client and read a dashboard. we bring - → A dedicated, single-tenant instance in your region for the duration. - → A named engineer on a shared channel throughout. - → Benchmark harnesses for the agreed criteria, with reproducible scripts. - → The report, and a sized configuration with a written quote. 02 · after the pilot ## After the pilot, the numbers decide. After the pilot, the same configuration continues at the published price — $570 / month for the default configuration, before add-ons — or you walk away with the report. ## Ready to measure? Send the workload. Tell us the data shapes, the volumes and the numbers you care about. We reply with a proposed scope. [Request a pilot](https://originchaindb.com/contact?topic=pilot) [See pricing](https://originchaindb.com/pricing) --- # Pricing | OriginChainDB Canonical source: https://originchaindb.com/pricing Sitemap last modified: 2026-09-27T14:44:27.000Z Pricing / One engine, four data models # Start free. Grow with your workload. Shared resources for your first idea. Dedicated resources when your application needs them. ## Free Explore For learning, prototypes, and your first application. $0/ month - 1 GB of database storage - Shared, scale-to-zero infrastructure - Query workbench and API access - One free database per account [Start free](https://app.originchaindb.com/signup) No card required. ## Configuration ·25 GB pooled Available A fixed monthly price for a smaller application. $29/ month - 25 GB database storage - 20 GB monthly data transfer - Shared compute · capped at 1 vCPU - Single node · India (Mumbai) [Choose 25 GB](https://app.originchaindb.com/instances) Monthly billing in USD. Review checkout before paying. ## Configuration ·50 GB pooled Available More storage and included transfer for your application. $49/ month - 50 GB database storage - 50 GB monthly data transfer - Shared compute · capped at 1 vCPU - Single node · India (Mumbai) [Choose 50 GB](https://app.originchaindb.com/instances) Monthly billing in USD. Review checkout before paying. ## Dedicated Configure Your own compute and storage, sized for your workload. From$320/ month - Single-tenant compute and storage - Configurable node size and capacity - Daily snapshots included - Optional standby and private networking [Estimate your setup](https://originchaindb.com/pricing#explore-3) Baseline estimate: 2 vCPU, 8 GB RAM, 20 GB storage, Mumbai. Amounts in USD, before applicable taxes. Final configuration and quote are confirmed before launch. [Need a custom deployment?](https://originchaindb.com/pricing#explore-4) What shapes the price ## A clear view of what you’re sizing. Adjust the resources behind your application. The estimator shows where each part of the total comes from. 01 ### Platform The base service and SQL. Vector, graph, and full-text have no separate model fee in this estimate. One engine 02 ### Resources Compute and storage per node, adjusted for the region and the size of your data. Size your workload 03 ### Recovery Optional standby and archive settings add resources to your configuration. Choose your recovery setup 04 ### Transfer Outbound data and optional private-network egress contribute to the total. Account for traffic Platform + resources + recovery + transferYour monthly estimate Dedicated / Configuration estimator ## Build a setup around your data. Start with compute, storage, and region. Open the workload and recovery options when you need them. ### Start with your capacity SQL included Region Mumbai is the launch baseline. Other regions are available on request. Compute per node4 vCPU · 16 GB Minimum nodes1 node Storage per node100 GB Estimated nodes and storage may increase to fit your data. Controls activate when the local calculator is ready. Size your dataOptional · vector, graph and full-text VectorsOff Dimensions Metadata per vector No vector workload selected. Graph nodesOff Relationships per node Properties per node No graph workload selected. Full-text documentsOff Average document size No full-text workload selected. Vector sizing estimates HNSW hot RAM. These are capacity assumptions, not measured query throughput. Recovery and networkingSingle zone · public subnet · recovery off by default Monthly writes100 GB Recovery windowOff Point-in-time recovery is in preview. Confirm retention and availability before launch. Monthly data transfer out100 GB Availability Standby replication is asynchronous. Cross-region cutover is operator-run. Networking ### Your configuration Planning estimate · USD / month · Taxes excluded —/month Awaiting calculation. JavaScript is needed to calculate this planning estimate. Defaults: 1 × 4 vCPU · 16 GB · 100 GB storage Estimated hot RAM— Add a workload to estimate its memory needs. Itemized costs appear when the calculator is ready. [Discuss your configuration](https://originchaindb.com/contact?topic=pricing) Final pricing, terms and available configurations are confirmed at launch. Compare the essentials ## Choose the resources. Keep building. [Explore the engine](https://originchaindb.com/product/database) Shared and dedicated database configurations | Your deployment | Free | 25 GB pooledAvailable | 50 GB pooledAvailable | Dedicated | | --- | --- | --- | --- | --- | | Monthly price · USD | $0 | $29 | $49 | From $320 · estimate | | Storage | 1 GB | 25 GB | 50 GB | Size it in the estimator | | Included data transfer | See free configuration limits | 20 GB / month | 50 GB / month | Set it in the estimator | | Compute | Shared | Shared · capped at 1 vCPU | Shared · capped at 1 vCPU | Single-tenant · configurable | | Nodes & availability | Shared infrastructure | Single node · no HA or replicas | Single node · no HA or replicas | Configurable nodes and standby | | Region | Assigned at launch | India (Mumbai) | India (Mumbai) | Mumbai; other regions on request | | Point-in-time recovery | Not included | Not included | Not included | Optional · in preview | | Hot-data memory | Fixed allocation | Not adjustable | Not adjustable | Depends on configuration | | How to start | Create an account | Choose in the console | Choose in the console | Review a quote before launch | Swipe to compare all configurations. The $49 pooled configuration is separate from the SQL Pro add-on. Check supported operations and service levels for your configuration. [Deployment guide ↗](https://originchaindb.com/docs/deploy)[Service levels ↗](https://originchaindb.com/sla)[Security ↗](https://originchaindb.com/security) Enterprise / Custom requirements ## Have a deployment to scope together? Bring your region, network, recovery, and support requirements. We’ll work through the configuration with you. [Discuss your deployment](https://originchaindb.com/contact?topic=enterprise) - 01 ### Private connectivity Review network boundaries and access requirements. - 02 ### Region & recovery Scope location, standby, and restore requirements. - 03 ### Your agreement Confirm deployment model, support, and commercial terms. Pricing questions ## Know what comes next. [Ask our team](https://originchaindb.com/contact?topic=pricing) What can I do with the free configuration? Start with a 1 GB database on shared infrastructure, explore the workbench, and connect through the API. It sleeps when idle. The console confirms eligibility before launch. [Create a free database](https://originchaindb.com/docs/free-account) What do the pooled configurations include? $29/month includes 25 GB storage and 20 GB monthly data transfer. $49/month includes 50 GB storage and 50 GB monthly data transfer. Both have shared compute capped at 1 vCPU, and a single node in India (Mumbai). Select a pooled configuration in the console and review checkout before paying. [Choose in the console](https://app.originchaindb.com/instances) Is the $49 pooled configuration the SQL Pro add-on? No. Configuration · 50 GB pooled is a database configuration. SQL Pro is a separate console add-on; it is not the product shown here. [Ask a pricing question](https://originchaindb.com/contact?topic=pricing) Are the four data models charged separately? The dedicated estimator includes SQL, vector, graph, and full-text without separate model fees. Their data and indexes still consume compute, memory, and storage; sizing those resources changes the estimate. [Understand the four models](https://originchaindb.com/product/database) Is the estimate my final bill? The estimator uses the website's bundled rate model. It is a planning estimate in USD, before applicable taxes. Review current availability, metered usage, and the final quote in the console or with our team before purchasing. [Discuss your configuration](https://originchaindb.com/contact?topic=pricing) What does the recovery option include? Dedicated instances include daily snapshots. Continuous archive and point-in-time recovery options are in preview. Availability, retention, and recovery behavior depend on the configuration you agree. [Review recovery behavior](https://originchaindb.com/docs/ops) Can I move between configurations? In-place changes between free, pooled, and dedicated configurations are not available yet. Pooled configurations do not support adding nodes, high availability, replicas, point-in-time recovery, or adjustable hot-data memory. [Discuss your requirements](https://originchaindb.com/contact?topic=pricing) What about annual billing or a custom agreement? Contact us to agree the term, support level, and any private deployment requirements. The estimator shows a monthly equivalent and does not apply an annual discount. [Talk to our team](https://originchaindb.com/contact?topic=enterprise) Your next step ## Start with your data. Start free, choose pooled capacity, or size a dedicated configuration. [Start free ↗](https://app.originchaindb.com/signup)[Talk to our team ↗](https://originchaindb.com/contact?topic=pricing) --- # Privacy policy - what OriginChain collects Canonical source: https://originchaindb.com/privacy Sitemap last modified: 2026-09-27T14:07:04.000Z privacy # Privacy notice OriginChain is a managed database service. The marketing site and the cloud service collect different things; both are set out below - what is collected, who it is shared with, where it is processed, how long it is kept, and the rights you have over it. last updated 27 September 2026 Browse this document - [The short version](https://originchaindb.com/privacy#tldr) - [Who we are](https://originchaindb.com/privacy#controller) - [What this website collects](https://originchaindb.com/privacy#website) - [OriginChain Cloud - the managed service](https://originchaindb.com/privacy#cloud) - [Categories of personal data](https://originchaindb.com/privacy#categories) - [Who we share data with](https://originchaindb.com/privacy#sharing) - [Where data is processed](https://originchaindb.com/privacy#transfers) - [How long we keep data](https://originchaindb.com/privacy#retention) - [Your rights as a Data Principal](https://originchaindb.com/privacy#rights) - [Consent and withdrawal](https://originchaindb.com/privacy#consent) - [How we protect data](https://originchaindb.com/privacy#security) - [Personal data breach](https://originchaindb.com/privacy#breach) - [Data Processing Addendum](https://originchaindb.com/privacy#dpa) - [Children's data](https://originchaindb.com/privacy#children) - [Changes to this notice](https://originchaindb.com/privacy#changes) ## The short version OriginChain is a managed database product operated by Silicoyn Technologies Pvt Ltd ("we"), an Indian private limited company headquartered in Bengaluru. You sign up, we run your managed instance for you. The rows you store are yours: we do not read them, sell them, or let them leave the region you picked. This notice is published in compliance with Section 5 of the Digital Personal Data Protection Act 2023 ("DPDP Act") and applies to all personal data we process, regardless of where you are located. ## Who we are Data Fiduciary: Silicoyn Technologies Pvt Ltd, a company incorporated under the Indian Companies Act 2013 with its registered office in Bengaluru, Karnataka, India. Grievance Officer / Data Protection Officer: reachable at grievance@originchain.ai. Complaints are acknowledged within 24 hours and resolved within 15 days, as required by Section 7(c) of the DPDP Act and the IT Rules 2021. ## What this website collects The marketing site is a static deployment. We do not set advertising cookies, we do not run third-party trackers, and we do not fingerprint visitors. Our content-delivery and hosting layer records standard access logs - IP address, user agent, referrer, timestamp - for abuse prevention and security monitoring. Those logs are retained for 30 days and then discarded. We do not link them to any other identifier. Lawful purpose: Section 4(2) DPDP Act - "specified purpose" of operating and securing a public website. ## OriginChain Cloud - the managed service When you sign up we collect the email, organisation name, and billing details you give us. Account-related personal data is stored in our managed account database in the Mumbai (Asia Pacific) region. That account system never touches the rows you put into your tenant. Your tenant lives in the region you pick. Rows, transaction logs, backups, and snapshots stay in that region. We do not replicate personal data across regions unless you explicitly request it. We collect operational metrics from your tenant - request rates, latencies, cache hit ratios, replication lag - over a private channel, and use them to run the service and show you dashboards. We never sample the content of your rows or your queries. Access to your tenant from our side is limited to a small on-call team, fully audit-logged, and only used to respond to a support ticket you file or to a live incident affecting your tenant. Lawful purpose: Section 4(2)(a) DPDP Act - performance of the service contract you signed up for; Section 4(2)(b) for tax / accounting record-keeping. ## Categories of personal data Identity & contact data: name, email address, organisation name. Billing data: tokenised card identifier (we never receive PAN/CVV - Razorpay does), billing address, GSTIN if provided, invoice history. Authentication data: hashed bearer tokens (Argon2id), session cookies, IP address of recent sign-ins. Service-content data: any rows you choose to store in your tenant. We treat the entire row blob as opaque - we do not classify, index for ML, or enrich it. Operational telemetry: request timestamps, error codes, latency buckets, cache-hit rates. No row content, no query content. ## Who we share data with Sub-processors used to deliver the Service: • Amazon Web Services, Inc. - cloud infrastructure: compute, storage, backups, DNS, object storage, our managed account database, and monitoring. • Razorpay Software Pvt Ltd - payment processing. Card data is tokenised; we never receive PAN or CVV. • Anthropic - natural-language compilation of /v1/ask requests when the LLM compiler runs. Prompts are submitted to the model and discarded; we do not store model output beyond the response we return. We do not sell personal data, share it with advertisers, or use it for cross-product marketing. We may disclose personal data when compelled by a valid Indian legal request (court order, summons under Section 91 CrPC, lawful direction of a government agency under Section 69 IT Act). We notify you of such requests where legally permitted. ## Where data is processed By default, all customer rows and tenant operational data stay in the Mumbai (Asia Pacific) region - Indian sovereign territory. Account-level data (email, billing) is processed in the Mumbai region. Cross-border transfer occurs only if (a) you select a non-Indian region for your tenant, or (b) you route /v1/ask traffic through a model endpoint in another region. Both are explicit choices on your part. Where the DPDP Act restricts cross-border transfer to specific countries, we comply with the restriction and notify you if your selected configuration becomes non-compliant. ## How long we keep data Tenant rows: as long as the database exists. When a database is deleted, its stored data is erased within 7 days and its backup snapshots within 30 days. Account / contact data: while the account is active, plus 7 years of billing record retention as required by Section 44AA of the Income Tax Act 1961. Authentication & security logs: 90 days. Marketing-site access logs: 30 days. ## Your rights as a Data Principal Under Sections 11–14 of the DPDP Act, you have the right to: (a) confirm whether we process your personal data and obtain a summary; (b) seek correction or erasure of inaccurate or unlawful data; (c) nominate a person to exercise these rights on your behalf in case of incapacity; (d) grievance redress for unsatisfactory responses. To exercise any right, email grievance@originchain.ai. We respond within 15 days. If unsatisfied, you may approach the Data Protection Board of India established under Section 18 of the DPDP Act. If you are an EU/UK resident and EU/UK data protection law applies, GDPR-equivalent rights apply: access, rectification, erasure, restriction, portability, objection. The same email reaches us; we respond within 30 days. ## Consent and withdrawal We rely primarily on contractual necessity (Section 4(2)(a) DPDP Act) - you signed up to use a service, we process the personal data needed to deliver it. For anything outside the strict service-delivery scope (e.g. opt-in marketing emails), we ask explicit consent and you can withdraw it at any time from /app/settings. ## How we protect data TLS 1.2+ on every external endpoint. Argon2id for stored credentials. AES-256 encryption at rest on all volumes and object storage (provider-managed keys; customer-managed keys available on Enterprise). Principle-of-least-privilege access control with no direct SSH - operator access is brokered and fully audit-logged. Daily encrypted backup snapshots into a 30-day-retention vault. See /security for the full posture. ## Personal data breach If we become aware of a personal data breach affecting you, we will notify you and the Data Protection Board of India in accordance with Section 8(6) of the DPDP Act and the breach-notification rules thereunder. The notification will describe the nature of the breach, categories and approximate number of affected records, likely consequences, and steps taken to mitigate it. ## Data Processing Addendum If you are processing personal data of others on the Service, you are the Data Fiduciary and Silicoyn Technologies Pvt Ltd is your Data Processor. The full DPA is at originchain.ai/dpa; email legal@originchain.ai for a counter-signed copy. ## Children's data OriginChain is not directed at children under 18 and we do not knowingly collect personal data of children. If you believe a child has provided us with personal data, email grievance@originchain.ai and we will delete it. ## Changes to this notice If this notice materially changes, we will update the date at the top of the page, note the change in our changelog, and email everyone on the managed service. We will not quietly start collecting anything new. ## Questions about any of this? Email us directly. We reply within a day. [support@originchain.ai](mailto:support@originchain.ai) --- # Distributed multimodal database | OriginChainDB Canonical source: https://originchaindb.com/product Sitemap last modified: 2026-09-24T06:54:16.000Z Distributed multimodal database # One database. More ways to build. SQL, vector, graph and full-text in one engine. Start with the model your application needs, then connect the others. [Start free](https://app.originchaindb.com/signup)[How it fits together](https://originchaindb.com/product/database) Four views of connected dataIllustrative example APPLICATION RECORDdoc-001One identity connects the representations you store. [SQLRead the record`title · author · category`Explore](https://originchaindb.com/product/sql)[VectorFind similar meaning`embedding → nearby records`Explore](https://originchaindb.com/product/vector)[GraphFollow a connection`document → author → topic`Explore](https://originchaindb.com/product/graph)[Full-textMatch the words`terms → ranked documents`Explore](https://originchaindb.com/product/fts) Manage access and recovery around the same database.[See how the models work together](https://originchaindb.com/product/database) ## Start with the foundation. Store your data, define access and choose a deployment that fits. ### Database Four data models in one engine, connected through shared record IDs. [Explore the database](https://originchaindb.com/product/database) ### Authentication Database users, roles and read policies shape access to your records. [Control access](https://originchaindb.com/product/authentication) ### Storage Choose dedicated capacity and plan recovery around regional snapshots. [Plan your storage](https://originchaindb.com/product/storage) ## Four ways to work with your data. Choose a query model for the question you need to answer. 01Records ### SQL Filter, join and aggregate the records behind your application. [Explore SQL](https://originchaindb.com/product/sql) 02Meaning ### Vector Search embeddings by similarity, then retrieve the records behind each match. [Explore vector search](https://originchaindb.com/product/vector) 03Connections ### Graph Declare relationships and follow the paths between connected records. [Explore graph](https://originchaindb.com/product/graph) 04Words ### Full-text Find terms and phrases in indexed text, with BM25 for ranked results. [Explore full-text search](https://originchaindb.com/product/fts) ## Or start with a question. Natural-language access translates questions into database queries. ### Natural language Ask a question about your tables and inspect the query plan behind the result. [Meet ASK](https://originchaindb.com/product/ask) ### Your model provider Connect your provider credentials for questions that need model-assisted compilation. [Explore provider keys](https://originchaindb.com/product/byok-llm) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Follow the quickstart](https://originchaindb.com/docs/quickstart) --- # Natural language — Ask questions. Get database results. | OriginChainDB Canonical source: https://originchaindb.com/product/ask Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) Natural language # Ask questions. Get database results. Turn a plain-English question into a query against your registered tables, with a plan you can inspect. [Start free](https://app.originchaindb.com/signup) [Read the ASK guide](https://originchaindb.com/docs/ask) ## From a question to a query. OriginChainDB turns your question into a query plan, runs it and returns database rows. Question and table scope, compiled into a query plan.Illustrative example Start with a question “Show the first 10 pending orders.” Table scope `shop.orders` ### Check the plan cache A matching cached plan can go straight to execution. Cache hit Reuse plan On a cache miss ### Try built-in query rules Recognized questions become a query plan. Only for questions outside those rules ### Ask the configured model The model helps compile a plan when available. Compiled plan Read`shop.orders` Filter`status = pending` Limit`10 rows` The database runs the plan order_idstatus order-042pending order-057pending Database rows, plus the plan when requested. 1. ### Give the question context Name the tables your question is about. Specific column names help OriginChainDB identify the data you need. 2. ### Turn the question into a query OriginChainDB tries built-in query rules first. Questions outside those rules can use the configured language model. The database then runs the resulting plan. Questions that need a model are sent to the configured provider. Keep sensitive values out of prompts. [Try an ASK example](https://originchaindb.com/docs/examples/ask) From your application ## Make the question concrete. One request carries the question, the tables to consider and whether to return the plan. `nl`The question Name the data and the condition you want. Here, we ask for up to ten pending orders. `schemas`The context Limit the compiler to relevant registered tables instead of your entire catalogue. `show_plan`The explanation Return the compiled plan with the rows so you can review the query before using it in your application. POST`/v1/tenants/:tenant/ask` ``` { "nl": "top 10 shop.orders where status = pending", "schemas": ["shop.orders"], "show_plan": true } ``` Example request · authenticated with your database credentials. [See a scoped request](https://originchaindb.com/docs/examples/ask/schemas-hint) Question → plan → review ## Make the answer easier to verify. Start with a specific question “Show important customers.” “Top 5 customers by total order amount.” Say which quantity matters and whether it is a total, count or individual value. Split unrelated requests into separate questions. 1. Check the selected tables.Confirm the plan reads the records you intended. 2. Review filters and limits.A valid plan can still express the wrong business question. 3. Test with familiar records.Compare returned rows with a small example you understand before adding the query to a workflow. [Inspect an example plan](https://originchaindb.com/docs/examples/ask/show-plan) ## Why ask with OriginChainDB? ### Focused context Choose the relevant tables so the compiler has less ambiguity to resolve in larger databases. [Choose relevant tables](https://originchaindb.com/docs/ask#ask) ### Visible execution Request the query plan and execution details alongside your results to understand how the question was answered. [Inspect a plan](https://originchaindb.com/docs/ask#plan) ### Reusable plans Matching questions can reuse compiled work. Cache status tells you whether a plan was reused. [Understand the cache](https://originchaindb.com/docs/ask#cache) ## Put it to work. ### Start with one table Try a focused question, inspect its plan, then build it into your application. [Follow the ASK guide](https://originchaindb.com/docs/ask) ### Choose your model setup Connect a provider key for questions that need a language model. [Explore provider keys](https://originchaindb.com/product/byok-llm) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Read the ASK guide](https://originchaindb.com/docs/ask) --- # Authentication & access control — Access starts with identity. | OriginChainDB Canonical source: https://originchaindb.com/product/authentication Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) Authentication & access control # Access starts with identity. Connect applications as database users, then define what those users can read. [Start free](https://app.originchaindb.com/signup) [Read the access guide](https://originchaindb.com/docs/dashboard/iam) ## Identify the caller. Shape the result. Roles, row policies and column masks determine which records and fields a database user can read. Read access for a restricted database userIllustrative example Authenticated database usersupport_7 Role Support reader Read request · customers 01 ### Stored records customer-01Assigned to support_7 customer-02Assigned to support_9 customer-03Assigned to support_7 The source table stays intact. ### Read policy 1. 1Check table access 2. 2Keep assigned rows 3. 3Mask private fields 02 ### Returned view customers2 eligible rows customer-01 emailMasked customer-03 emailMasked Fields filtered for this caller A restricted view of the same data. 1. ### Choose the credential A database-user key identifies the caller so requests can use that user's roles and grants. 2. ### Define the view Grant table access, filter eligible rows and mask sensitive columns. A support role could read assigned customers while their private fields remain hidden. Use a service-enforced read-only token when an application must not write. [Explore row policies](https://originchaindb.com/docs/dashboard/iam#rls) ## Why OriginChainDB ### Permissions close to the data Define database roles and table access alongside the records they protect. [Roles and grants](https://originchaindb.com/docs/dashboard/rbac#database-rbac) ### Precise read access Filter rows and mask columns without maintaining a separate copy of each table for every role. [Policies and masking](https://originchaindb.com/docs/dashboard/iam#rls) ### Credentials with ownership Rotate or revoke a personal token independently of the shared database bearer. [Credential lifecycles](https://originchaindb.com/docs/dashboard/iam#credentials) Design the access model ## One table. Different responsibilities. Start with the job each caller needs to do. Give support a view of assigned customers; give an analyst the records needed for reporting. This example combines table read access, row filtering and field masking. The roles are illustrative, not built-in defaults. [Build a row policy](https://originchaindb.com/docs/dashboard/iam#rls) Example · customer read permissions | Role | Rows returned | Email field | | --- | --- | --- | | Support reader | Assigned customers | Masked | | Regional analyst | Customers in their region | Masked | | Contact manager | Assigned customers | Visible | Check the result as the intended database user. An administrative bearer is not a test of restricted access. Credentials over time ## Know who connects. Know how to revoke access. Keep each credential tied to its purpose and owner, from the first connection through offboarding. 1. 01 ### Issue Choose the identity and credential type for the application or person. 2. 02 ### Store Save a newly issued secret in your secret manager when it is shown. 3. 03 ### Rotate or revoke Update clients using a shared bearer. Personal-token changes affect that token independently. Team membership ≠ database identity Removing an organization member does not itself delete their personal tokens. Review database users and shared credentials as part of offboarding. [Read the offboarding guidance](https://originchaindb.com/docs/dashboard/rbac#guidance) ## Put it to work. ### Separate team and database access Organization membership and database identity are separate. Adding someone to your team does not grant permission to query records. [Understand identities](https://originchaindb.com/docs/dashboard/iam#database-users) ### Verify access in your application Check how each query path applies the user's roles and policies. Recorded write grants alone do not prevent writes. [Review enforcement](https://originchaindb.com/docs/dashboard/iam#enforced) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Read the access guide](https://originchaindb.com/docs/dashboard/iam) --- # Model provider keys — Your model. Your provider key. | OriginChainDB Canonical source: https://originchaindb.com/product/byok-llm Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) Model provider keys # Your model. Your provider key. Choose the provider and model that help turn natural-language questions into database queries. [Start free](https://app.originchaindb.com/signup) [Configure a provider](https://app.originchain.ai/settings/integrations/llm) ## Connect a model to ASK. Save model settings for your instance, then use the same ASK endpoint from your application. Instance configuration connects the ASK compiler to your selected provider.Illustrative example Your OriginChainDB instanceSaved configuration ### Provider key •••• •••• •••• Encrypted in storage Used on the server Your application’s question /askQuery compiler Rules and cached plans can avoid a model call. The dashboard displays a masked key. Only when a model is neededQuestion + model request Selected provider ### Your chosen model Helps translate the question into a query plan. Plan returned to the engine Configuration belongs to the instance.Provider terms govern model processing. 1. ### Choose your provider Select an available provider and model, add your key and save the settings for your instance. Connection options depend on the selected provider. 2. ### Use it when needed ASK uses the enabled provider when a question needs model assistance. Built-in query rules and cached plans can avoid a model call. Model availability depends on the provider and instance configuration. Disabling BYOK returns ASK to the engine's default configuration. [See how ASK works](https://originchaindb.com/docs/ask) ## Why connect your provider? ### Model choice Update the saved configuration as your application's requirements change, while keeping the same ASK request interface. [Open provider settings](https://app.originchain.ai/settings/integrations/llm) ### Managed credentials Your key is encrypted before it is saved. The dashboard shows only a masked value. [Manage your key](https://app.originchain.ai/settings/integrations/llm) ### Inspectable queries Review the resulting plan alongside database rows to understand how the question was executed. [Inspect a query plan](https://originchaindb.com/docs/ask#plan) A deliberate setup ## Choose. Connect. Verify. [Open provider settings](https://app.originchain.ai/settings/integrations/llm) Decide which model should handle questions that need assistance, then verify it with a small query you can check. 1. ### Choose an available model Review the provider’s supported models, account access and data-processing terms. Use a model available in your instance’s provider catalogue. 2. ### Save the configuration Add the provider key in instance settings. Supply any endpoint options the selected provider requires. Keep credentials out of application source code. 3. ### Review a first result Send a focused ASK request with plan inspection enabled. Check its tables, filters and returned rows before relying on it in a workflow. Keep control as requirements change ## Manage the key. Understand the next request. Provider credentials and database credentials serve different purposes. Changing your model key changes the model connection; your application still uses ASK. [Review the ASK contract](https://originchaindb.com/docs/ask) Replace the saved key Subsequent requests use the updated provider configuration. Disable BYOK Keep the saved configuration, while ASK uses the engine’s default model settings. Remove the configuration Delete the saved connection and return to engine defaults. This does not revoke the key at the provider. Check model charges and limits with your provider. Database pricing is separate. ## Put it to work. ### Build with ASK Scope questions to registered tables and request a plan when reviewing query behaviour. [Read the ASK guide](https://originchaindb.com/docs/ask) ### Plan your costs Provider terms apply when using your credentials. Review database pricing separately from model usage. [Review database pricing](https://originchaindb.com/pricing) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Configure a provider](https://app.originchain.ai/settings/integrations/llm) --- # Four data models. One database. | OriginChainDB Canonical source: https://originchaindb.com/product/database Sitemap last modified: 2026-09-24T06:54:16.000Z Distributed multimodal database # Four data models. One database. Query records, vectors, relationships and text in one engine. [Start free](https://app.originchaindb.com/signup) [Read the quickstart](https://originchaindb.com/docs/quickstart) ## One record. Four ways to use it. Keep business fields, embeddings, searchable text and relationships in one engine, linked by record ID. [Explore the architecture](https://originchaindb.com/architecture) One record, four representationsIllustrative example FOLLOW A RECORD SHARED RECORD ID `prod-042` Alpine shell jacket commerce.products A product your application can retrieve in four ways. OriginChainDB · one engine 01 / SQL[↗](https://originchaindb.com/product/sql) ### Read the facts. id prod-042 category outdoor in_stock true Filter and query business fields. 02 / VECTOR[↗](https://originchaindb.com/product/vector) ### Find the meaning. “A lightweight layer for wet trails” 03 / GRAPH[↗](https://originchaindb.com/product/graph) ### Follow the context. `prod-042`belongs_toOutdoor gear Traverse declared record relationships. 04 / FULL-TEXT[↗](https://originchaindb.com/product/fts) ### Match the words. waterproofjacket `prod-042`term match Search the content you index. Linked by ID. Queried through the model that fits.Example representations · no live queries 1. ### Model and index Define fields and graph relations. Write rows, then add your embeddings and search entries using the same record ID. 2. ### Query your way Filter with SQL, find similar meaning with vectors, match keywords with full-text search and follow connections with graph queries. The current API uses separate row, vector and full-text calls; each call is atomic. [See an ingestion example](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog) An application workflow ## From a search to a useful result. Product discovery can use all four models. Your application coordinates the separate requests and decides what to return. Example customer search “Lightweight shoes for the trail” Meaning + keywords + business context [Explore retrieval and ranking](https://originchaindb.com/docs/rag#retrieve) 1. VectorFull-text ### Find candidate products Search embeddings for similar meaning and indexed text for matching words. Each query returns candidate record IDs. 2. Your application ### Combine the candidates Merge and deduplicate those IDs in your application. Choose how the two ranked lists contribute to the result. 3. SQL ### Apply business rules Fetch the candidate rows with SQL, filtering for available stock and the customer's price range. 4. Graph ### Add relationship context Follow each product's declared supplier relation, then assemble the product and supplier details in your application. Before your first integration ## Make three decisions early. How will models connect? Choose a stable primary key and reuse it for vector IDs and full-text document IDs. [Define the schema](https://originchaindb.com/docs/schemas) Which embeddings will you use? Set the model, dimensions and distance metric. Store the metadata needed for vector filtering. [Plan vector data](https://originchaindb.com/docs/schemas/vector) What happens when data changes? Handle row, vector and text-index writes separately. Update affected indexes and track calls that need retrying. [Follow the ingestion example](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog) ## Why OriginChainDB? Bring application data and AI retrieval into the same database. ### Less integration work Keep records, search indexes and relationships in one database, reducing the services and synchronization pipelines you maintain. [Inside the engine](https://originchaindb.com/architecture) ### Richer retrieval Combine semantic matches, exact keywords and relationship context in your application for RAG, search and recommendations. [Build a retrieval flow](https://originchaindb.com/docs/rag) ### Results you can inspect Review published throughput, recall and retrieval-quality measurements, with the hardware, datasets and commands used to produce them. [View the benchmarks](https://originchaindb.com/benchmarks) ## Start small. Scale with dedicated resources. [Deployment guide](https://originchaindb.com/docs/deploy) ### Free Start on shared infrastructure. Build your first application and explore the engine. [Explore the free plan](https://originchaindb.com/docs/free-account) ### Dedicated Choose your region and dedicated resources. Configure compute and storage for your workload. [Configure your instance](https://originchaindb.com/pricing) Optional standby replication is asynchronous. Choose your recovery requirements alongside compute and storage. [Understand the topology](https://originchaindb.com/docs/deploy#replication) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup) [Read the quickstart](https://originchaindb.com/docs/quickstart) --- # Full-text search — Find the words that matter. | OriginChainDB Canonical source: https://originchaindb.com/product/fts Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) Full-text search # Find the words that matter. Search document text with keyword and phrase matching, or rank relevant results with BM25. [Start free](https://app.originchaindb.com/signup) [Build your first search](https://originchaindb.com/docs/fts) Inside full-text search: index, match and rankIllustrative example THE QUERYvector searchWords become a route into your content. 01Index the content doc-12title + body A guide to vector search Use embeddings to find relevant content. Vector search compares meaning. doc-27title + body Choosing a search index 02Find candidate documents vector doc-12doc-41 search doc-12doc-27doc-41 BM25 weighsTerm frequencyTerm rarityDocument length 03Rank useful matches 1. 01 A guide to vector searchdoc-12Both terms in the title and body 2. 02 Choosing a search indexdoc-27Related text in the body 3. 03 Working with embeddingsdoc-41A shorter mention of vector search Example ordering, not measured relevance scores. ## Index text. Retrieve its record. Make descriptions, articles and application content searchable. A shared document ID connects each search hit to its original record. 1. ### Make text searchable Write a search entry under a document ID that matches your record's primary key. Choose the table and field you will query later. 2. ### Choose the query Match terms or phrases, use BM25 for ranked results and fetch the corresponding records by ID when you need their full context. Update the search entry when its source text changes. Row edits do not automatically re-index text. [Index a document](https://originchaindb.com/docs/examples/fts/index-doc) ## Why search on OriginChainDB? ### Relevance you can inspect Use BM25 scoring and its explain output to understand why documents match and how their terms contribute to a result. [Explore BM25 ranking](https://originchaindb.com/docs/examples/fts/bm25) ### More precise queries Use the JSON query DSL for field weights, filters, negation and multi-field search when a simple keyword query is not enough. [Explore search queries](https://originchaindb.com/docs/fts) ### Connected application data Keep searchable text alongside SQL, vector and graph data, with shared IDs linking search results to the rest of your application. [See the data models](https://originchaindb.com/product/database) QUERY DESIGN ## Three questions. Three ways to search. Choose the simplest search mode that answers your application's question. ### Boolean Does this document match? Include or exclude content by its terms when relevance ordering is not needed. [See an example](https://originchaindb.com/docs/fts) ### Phrase Do these words appear together? Match an exact sequence such as a product name, quoted passage or technical phrase. [See an example](https://originchaindb.com/docs/examples/fts/phrase) ### BM25 Which matches are most relevant? Rank candidates using term frequency, term rarity and document length. [See an example](https://originchaindb.com/docs/examples/fts/bm25) BM25, phrase search and advanced query options require the corresponding full-text capability on your instance. RELEVANCE THAT FITS YOUR CONTENT ## Give your important fields more say. A match in a product title may matter more than a passing mention in its description. Use multi-field weights to express that choice, then inspect the results. [Try a weighted query](https://originchaindb.com/docs/examples/fts/multi-field) Refine the resultAfter indexing Field weights Choose the relative importance of each searchable field. Highlights & facets Help users see where a match appears and explore grouped results with supported BM25 options. Explain Review why a document scored within this query. Scores are not comparable across different queries. [Read the search reference](https://originchaindb.com/docs/fts) ## Put it to work. ### Start with ranked search Index text and return relevant matches using BM25. The guide covers required capabilities and response formats. [BM25 example](https://originchaindb.com/docs/examples/fts/bm25) ### Connect an Elasticsearch client Use supported Elasticsearch 7.x clients through the compatibility API, with documented search operations, aggregations and limits. [Elasticsearch connection guide](https://originchaindb.com/docs/connect-elasticsearch) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Build your first search](https://originchaindb.com/docs/fts) --- # Graph — Your data has connections. Follow them. | OriginChainDB Canonical source: https://originchaindb.com/product/graph Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) Graph # Your data has connections. Follow them. Model relationships between records, traverse them with Cypher or the graph API, and explore the paths that connect them. [Start free](https://app.originchaindb.com/signup) [Build your first relationship](https://originchaindb.com/docs/graph/quickstart) Customers, orders and products connected by declared relations.Illustrative example shop namespace7 records · 6 relationships Customers Orders Products placed_by contains Customercustomer-7Maya Customercustomer-8Leo Orderorder-42Start here Orderorder-43Paid Orderorder-44Pending Productproduct-9Trail runner Productproduct-12Day pack - Start hereorder-42 placed_byMayacustomer-7 containsTrail runnerproduct-9 - Orderorder-43 placed_byMayacustomer-7 containsDay packproduct-12 - Orderorder-44 placed_byLeocustomer-8 containsTrail runnerproduct-9 Selected order's references Other declared relationships ## Turn references into relationships. An order belongs to a customer. A document covers a topic. Declare those connections and query them alongside the records they describe. 1. ### Declare a relation Connect a source field to another table's primary key. Row writes maintain the edge indexes for that declared relationship. 2. ### Explore the connections Start from a record, follow bounded hops and filter the records you reach. Use Cypher or choose a dedicated graph endpoint. Cypher supports a defined subset. Start traversals with a record's primary key; some algorithms require relations within one table. [Try a bounded traversal](https://originchaindb.com/docs/examples/graph/bfs) Model the connection ## The relationship starts in your record. A customer's ID on an order is already a connection. Give that field a named relation so the engine can follow it. Source fieldshop.orders.customer_id The value written on each order. Relation nameplaced_by The name used in a graph query. Target keyshop.customers.id The customer the order points to. shop.ordersRelation block · TOML ``` [[relations]] name = "placed_by" from_col = "customer_id" bidirectional = true [relations.target] namespace = "shop" table = "customers" pk = "id" ``` Register the source and target tables, then add this relation to the order schema. Row writes maintain its edges. Enable both directions when you also need to find a customer's orders. [Build a graph schema](https://originchaindb.com/docs/schemas/graph) Ask a connected question ## Choose the method that fits the question. A nearby record, a route and an influential node are different questions. Start with the smallest traversal that answers yours. 01 ### What is connected? Neighbours and reverse lookups Follow a declared relation in either configured direction to find adjacent records. 02 ### How do we get there? Bounded traversal and paths Explore a neighbourhood with BFS, or choose a path method for the route you need. Set a depth bound. 03 ### What patterns emerge? Ranking and communities Run PageRank over a nominated set of nodes or evaluate community methods. Check each algorithm's relation requirements first. Start with one order ### Retrieve its customer. Pin the starting record and return the properties your application needs. [Explore graph queries](https://originchaindb.com/docs/graph) ``` MATCH (o:orders {id: 'order-42'}) -[:placed_by]->(c) RETURN c.id, c.name; ``` ## Why graph on OriginChainDB? ### Relationships beside records Keep business fields and graph relations in one engine, using the same record identities throughout your application. [Define a graph schema](https://originchaindb.com/docs/schemas/graph) ### More ways to investigate Find neighbours and paths, then explore ranking or community algorithms for the questions your relation can answer. [Explore graph methods](https://originchaindb.com/docs/graph) ### Context after retrieval Follow related people, topics or products after a search result to give your application more useful context. [Explore retrieval patterns](https://originchaindb.com/docs/rag) ## Put it to work. ### Model your first connection Declare source fields, target tables and traversal direction before writing records that form your graph. [Graph schema guide](https://originchaindb.com/docs/schemas/graph) ### Choose a graph method Compare traversal, path, ranking and community endpoints, with request examples and the relation requirements for each. [Graph API reference](https://originchaindb.com/docs/graph) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Build your first relationship](https://originchaindb.com/docs/graph/quickstart) --- # SQL — Your application data. Ready to query. | OriginChainDB Canonical source: https://originchaindb.com/product/sql Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) SQL # Your application data. Ready to query. Filter, join and aggregate records in the same database as your vectors, relationships and search indexes. [Start free](https://app.originchaindb.com/signup) [Run your first SQL query](https://originchaindb.com/docs/sql/quickstart) ## From question to result. Use SQL for the questions behind your application: which orders need attention, what customers bought and how activity changes over time. Related tables. One query. A regional revenue report.Illustrative example 01 Related tables shop.customers | id | region | | --- | --- | | c-01 | IN | | c-02 | EU | shop.orders | Customer | Cents | Status | | --- | --- | --- | | c-01 | 2,400 | paid | | c-02 | 3,600 | paid | | c-01 | 1,200 | paid | | c-02 | 900 | pending | 02 One query 1. JOINMatch customers`customer_id = id` 2. WHEREKeep paid orders`status = 'paid'` 3. GROUP BYTotal by region`SUM(total_cents)` 03 Application result Paid revenueBy customer region Illustrative query result, revenue in cents | Region | Revenue | | --- | --- | | IN | 3,600 | | EU | 3,600 | Pending orders stay out of the total. All amounts shown in cents. 1. ### Define your tables Declare fields, primary keys and indexes. Write records through the SQL or row API, keeping related tables in the namespaces your application uses. 2. ### Run and inspect Send a SQL statement, read the returned rows and use EXPLAIN to check how the engine scans, filters and joins your data. OriginChainDB implements a documented SQL subset. Supported statements differ between the HTTP API and PostgreSQL wire access. [Try a join](https://originchaindb.com/docs/examples/sql/left-join) In your application ## A lookup. A join. A clearer picture. Start with the records behind a screen. Join related tables when the question crosses them, then aggregate when you need a summary. The revenue query uses the sample tables above. It matches each paid order to a customer and adds the amounts within each region. [Read supported SQL syntax](https://originchaindb.com/docs/sql) Query patternsSQL 01Join and aggregate Paid revenue by region ``` SELECT c.region, SUM(o.total_cents) AS revenue_cents FROM shop.orders o JOIN shop.customers c ON o.customer_id = c.id WHERE o.status = 'paid' GROUP BY c.region; ``` 02Find the records you need The latest paid orders ``` SELECT id, customer_id, total_cents FROM shop.orders WHERE status = 'paid' ORDER BY placed_ms DESC LIMIT 20; ``` 03Bind application input Pass a customer ID as $1 ``` SELECT id, total_cents FROM shop.orders WHERE customer_id = $1 LIMIT 20; ``` Send the value for `$1` in the request's `params` array. Understand the work ## Inspect before you tune. A query plan shows how an answer is assembled. Use it to make a deliberate change, then check what that change achieved. 1. 01`EXPLAIN` ### Read the plan Check whether the engine scans the table or uses an index. Follow the filters, joins and aggregation stages. 2. 02`WHERE + index` ### Narrow the input Filter on an indexed field where possible. Return the columns you need and add a limit to exploratory queries. 3. 03`EXPLAIN ANALYZE` ### Measure the execution Run the query with execution details, compare the work performed and keep changes that help your workload. Keep the whole query in view. Sorting can process the full input before a limit applies. Narrow that input first when working with large tables. [Explore query plans](https://originchaindb.com/docs/dashboard/explain) ## Why SQL on OriginChainDB? ### Familiar query patterns Build application screens and reports with filtering, sorting, joins and aggregates. Bind user input through positional parameters. [Explore the SQL reference](https://originchaindb.com/docs/sql) ### Records behind every result Use shared record IDs to retrieve the business fields behind a vector match, graph connection or full-text search hit. [See the data models](https://originchaindb.com/product/database) ### Execution you can inspect Review query plans and index use before tuning a query, with EXPLAIN available through the same SQL endpoint. [Inspect a query](https://originchaindb.com/docs/dashboard/explain) ## Put it to work. ### Find the right SQL pattern Check supported statements, response formats and query limits before adding a new query to your application. [SQL reference](https://originchaindb.com/docs/sql) ### Connect a database client Use an enabled PostgreSQL wire endpoint. Check its preview scope and the read-only rules for database-user sessions. [Client connection guide](https://originchaindb.com/docs/connect-sql-client) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Run your first SQL query](https://originchaindb.com/docs/sql/quickstart) --- # Database storage — A home for your data. A plan for recovery. | OriginChainDB Canonical source: https://originchaindb.com/product/storage Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) Database storage # A home for your data. A plan for recovery. Store application data with dedicated capacity, regional snapshots and documented recovery options. [Start free](https://app.originchaindb.com/signup) [Review deployment options](https://originchaindb.com/docs/deploy) ## Choose capacity. Prepare recovery. Daily storage snapshots provide the recovery baseline for dedicated configurations. Regional storage, snapshots and recoveryIllustrative example Your chosen regionDedicated configuration WriterActive application data Asynchronous replicationStandby may lag Optional standbyAvailability configuration Daily storage snapshotRecovery baseline Encrypted snapshot archiveStored in the instance's region Recover into a fresh instanceChoose an available recovery point 1. Earlier snapshotRetained history 2. Selected snapshotRestore source 3. Later changeFor example, a bad import 4. Fresh instanceValidate, then reconnect Optional archived changes extend a snapshot when available. Archive streaming remains a preview. 1. ### Configure your instance Choose a region and dedicated compute and storage for your workload. Review capacity, availability and recovery settings together. 2. ### Recover from an available point Restore into a fresh instance using the selected snapshot and available archived changes. Check the recoverable point before choosing your target. Archive streaming is an opt-in preview; plan current recovery around daily snapshots. [How recovery works](https://originchaindb.com/docs/ops#backups) ## Why OriginChainDB ### Dedicated resources On dedicated configurations, database compute and storage are separate from other customers. [Configure capacity](https://originchaindb.com/docs/dashboard/create-instance#dedicated) ### Regional backups Dedicated storage snapshots are encrypted and archived in the instance's region. [Snapshot details](https://originchaindb.com/docs/ops#backups) ### Clear recovery behavior Review write durability, archive status and operational limits before setting your recovery expectations. [Durability and recovery](https://originchaindb.com/docs/ops#durability) Choose a deployment ## Capacity and protection are separate decisions. Start with the workload, then choose the resources and recovery options it needs. Storage choices by configuration | What you choose | Free | Dedicated | | --- | --- | --- | | Resources | Shared starting configuration | Dedicated compute and storage | | Capacity | Current free-plan quota | Choose capacity for the workload | | Recovery | Point-in-time recovery not included | Daily snapshots; review available recovery points | | Availability | Review the plan's limits | Optional asynchronous standby | Scroll to compare both plans A standby supports availability. Snapshots provide recovery history. Choose each for the failure you need to handle. [Configure a dedicated instance](https://originchaindb.com/docs/dashboard/create-instance#dedicated) A practical restore plan ## Recover the data. Verify the application. A successful restore is more than a running instance. Check that the recovered records and your application's connections match the state you expect. [Read the recovery runbook](https://originchaindb.com/docs/ops#backups) 1. 01 ### Choose an available point Inspect the snapshot and archive state. Select a recoverable point before the unwanted change. 2. 02 ### Restore a fresh instance Recovery uses the chosen snapshot and available archived changes. Keep the source available for comparison. 3. 03 ### Validate before switching Check application-critical records, query results and permissions. Then update the application to use the recovered instance. Plan against the recoverable point. Archive streaming is still a preview. Current recovery planning should use daily snapshots, rather than a promised per-second recovery window. [Check archive limitations](https://originchaindb.com/docs/ops#continuous-backup-streaming) ## Put it to work. ### Start with the right capacity Free offers a shared starting configuration. Dedicated opens resource and recovery choices; free configurations do not include point-in-time recovery. [Compare deployment choices](https://originchaindb.com/docs/dashboard/create-instance#choose) ### Plan backups and availability A standby and a backup serve different purposes. Standby replication is asynchronous; recent acknowledged writes may be lost if the writer fails. [Understand standby recovery](https://originchaindb.com/docs/ops#failover) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Review deployment options](https://originchaindb.com/docs/deploy) --- # Vector search — Find similar meaning in your data. | OriginChainDB Canonical source: https://originchaindb.com/product/vector Sitemap last modified: 2026-09-23T08:33:23.000Z [All products](https://originchaindb.com/product) Vector search # Find similar meaning in your data. Store embeddings, retrieve nearby matches and connect each result to its application record. [Start free](https://app.originchaindb.com/signup) [Run your first similarity query](https://originchaindb.com/docs/vector/quickstart) Explore how a question finds nearby meaningIllustrative example EXPLORE A QUERY Meaning has neighbours.Conceptual embedding space Query vector Similar candidates A simplified projection. Search uses the full embedding, not this picture. YOUR QUESTION “Something warm for a mountain hike” Similarity search → record IDs 1. 01Insulated trail jacket 2. 02Lightweight fleece layer 3. 03All-weather hiking shell Illustrative matches · no live query ## Embed your data. Search by similarity. Find relevant content even when a query uses different words. Your application supplies the embeddings; OriginChainDB indexes and searches them. 1. ### Store your embeddings Generate vectors with your model and write them with record IDs. Include the metadata you need for filtering at write time. 2. ### Retrieve useful matches Submit a query embedding, choose a distance metric and fetch the records behind the returned IDs. Keep model and dimensions consistent. Metadata filters use exact equality. Filtering behaviour varies by index; IVF-PQ filtering is not supported. [Try a metadata filter](https://originchaindb.com/docs/examples/vector/topk-filtered) CHOOSE THE TRADE-OFF ## Your workload. Your index. Memory, query effort and recall pull in different directions. Choose with measurements from your own dataset. HNSW ### Navigate a neighbour graph A useful starting point for similarity search. Tune the graph and search breadth, then measure the recall your application needs. Metadata filtering: candidate post-filtering IVF ### Search selected partitions Organize vectors into cells and control how many are probed. More probes inspect more candidates, with a corresponding query cost. Metadata filtering: during cell scans IVF-PQ ### Compress the representation Evaluate a smaller vector representation when memory matters. Test the effect of compression on your data and recall target. Metadata filtering: not supported [Compare index configuration](https://originchaindb.com/docs/vector) Dense rankingA · C · BEmbedding similarity Sparse rankingB · A · DSparse-vector similarity Reciprocal rank fusionA · B · C · DIllustrative ordering · equal weights HYBRID VECTOR RETRIEVAL ## Two representations. One ranked list. Write dense and sparse vectors under matching IDs. The hybrid endpoint combines their rankings so your application can retrieve one set of candidates. This endpoint fuses vector rankings. Full-text BM25 is a separate search path. [Build a hybrid query](https://originchaindb.com/docs/examples/vector/hybrid-topk) ## Why vectors on OriginChainDB? ### Context beside similarity Use shared IDs across records, vectors, graph relationships and searchable text to assemble the context your application needs. [Explore retrieval patterns](https://originchaindb.com/docs/rag) ### Index choices Start with HNSW, then evaluate IVF or IVF-PQ when your workload calls for a different memory and recall trade-off. [Compare vector indexes](https://originchaindb.com/docs/vector) ### Dense and sparse together Use the hybrid vector endpoint to combine both rankings with reciprocal rank fusion and return one list of matches. [Try hybrid retrieval](https://originchaindb.com/docs/examples/vector/hybrid-topk) ## Put it to work. ### Choose an index for your workload Review index setup, filtering behaviour and memory requirements. Measure recall on your own data before changing the default. [Vector index guide](https://originchaindb.com/docs/vector) ### Combine dense and sparse vectors Write both representations under the same ID and fuse their ranked results through the hybrid top-k endpoint. [Hybrid retrieval example](https://originchaindb.com/docs/examples/vector/hybrid-topk) ## Build with OriginChainDB. [Start free](https://app.originchaindb.com/signup)[Run your first similarity query](https://originchaindb.com/docs/vector/quickstart) --- # Cancellation & Refund Policy | OriginChainDB Canonical source: https://originchaindb.com/refund-policy Sitemap last modified: 2026-09-27T14:07:04.000Z Policies & agreements # Cancellation & Refund Policy This policy explains how to cancel and when we refund payments for OriginChainDB, which is operated by Silicoyn Technologies Pvt Ltd, Bengaluru, India. It forms part of our [Terms of Service](https://originchaindb.com/terms). Last updated 27 September 2026 Browse this document - [Billing](https://originchaindb.com/refund-policy#billing) - [Cancelling](https://originchaindb.com/refund-policy#cancelling) - [Deleting your account](https://originchaindb.com/refund-policy#account) - [Refunds](https://originchaindb.com/refund-policy#refunds) - [Requesting a refund](https://originchaindb.com/refund-policy#request) - [Missed card payments](https://originchaindb.com/refund-policy#missed-payments) - [Contact](https://originchaindb.com/refund-policy#contact) ## Billing - Paid configurations are monthly subscriptions, billed in US dollars. - You pay for the first month when you place your order, then on the same day each month until you cancel. - Pooled configurations are paid by card. Dedicated configurations are confirmed by our team before they start, and are billed as set out in your order. - Payments are processed by Razorpay. ## Cancelling - You can cancel at any time by deleting the configuration in the console. - Deleting stops the service and permanently deletes that database straight away; this cannot be undone, so export any data you need first. Its stored data is erased within 7 days and its backups within 30 days. - Deleting also cancels that configuration's subscription, so you won't be charged for it again. ## Deleting your account - Card billing is cancelled straight away, so you won't be charged again. - Each pooled configuration keeps running until the end of the period you've already paid for, and is then deleted. - You have 30 days to change your mind: sign back in and your account is restored. Card billing is not restarted, so a pooled configuration is still deleted at the end of its paid period; order it again if you need it. - During those 30 days a database may be suspended; we email you before that happens. - After 30 days, every database left in your account is deleted, and its stored data is erased within 7 days. ## Refunds - Payments are non-refundable, including for a partly used month, except in these cases, which we refund in full: - you were charged by mistake, for example twice for the same month; - you were charged after you cancelled; - your database could not be created and never became available; - the law requires a refund. - Credits for availability shortfalls are covered by our [SLA](https://originchaindb.com/sla), as described in the [Terms of Service](https://originchaindb.com/terms). ## Requesting a refund - Email [support@originchain.ai](mailto:support@originchain.ai) within 30 days of the charge, with your account email, the configuration name and the payment ID from your receipt. - We acknowledge requests within 24 hours and resolve them within 15 days. - Approved refunds go back to the card you paid with; your bank usually shows them within 5-10 business days. ## Missed card payments - If a monthly payment fails, it is retried automatically. - If it still can't be collected, we email you, and your configuration keeps running for 7 more days before it is suspended. - Suspension keeps your data. Once the payment is made, the configuration is restored. Email [support@originchain.ai](mailto:support@originchain.ai) if you need help paying. ## Contact Support: [support@originchain.ai](mailto:support@originchain.ai) · Grievances: [grievance@originchain.ai](mailto:grievance@originchain.ai) · Silicoyn Technologies Pvt Ltd, Bengaluru, Karnataka, India --- # OriginChainDB developer resources | OriginChainDB Canonical source: https://originchaindb.com/resources Sitemap last modified: 2026-09-28T07:29:43.000Z The developer library # From the first query to your next application. Guides, reference material and engineering notes for building with OriginChainDB. [Open the docs](https://originchaindb.com/docs)[Explore the architecture](https://originchaindb.com/architecture) ## A path for the task in front of you. Start anywhere. Each step takes you to the source material. Choose your starting pointBuild / Evaluate / Follow 01 ### Build a first workflow Go from a connection to a query you can inspect. 1. [ConnectConfigure your endpoint](https://originchaindb.com/docs/quickstart) 2. [ModelDefine fields and relationships](https://originchaindb.com/docs/schemas) 3. [QueryChoose a query model](https://originchaindb.com/docs/query) 02 ### Evaluate the database Understand the design and test your workload. 1. [ArchitectureFollow the write and read paths](https://originchaindb.com/architecture) 2. [MeasurementsReview setup and methodology](https://originchaindb.com/benchmarks) 3. [SecurityReview controls and responsibilities](https://originchaindb.com/security) 03 ### Follow the engineering Read the decisions behind the implementation. 1. [Release notesSee what changed](https://originchaindb.com/changelog) 2. [NewsroomCompany news and industry perspectives](https://originchaindb.com/newsroom) 3. [Engineering blogGo deeper on a design](https://originchaindb.com/blogs) 4. [CommunityJoin the conversation](https://community.originchain.ai/) Working examples ## Go straight to a query. [All examples](https://originchaindb.com/docs/examples) [SQL ### Which records match? Open the guide](https://originchaindb.com/docs/sql/quickstart)[Vector ### What is similar? Open the guide](https://originchaindb.com/docs/vector/quickstart)[Graph ### How are they connected? Open the guide](https://originchaindb.com/docs/graph/quickstart)[Full-text ### Where do these words appear? Open the guide](https://originchaindb.com/docs/fts) [HTTP API reference](https://originchaindb.com/docs/api)[Client SDKs](https://originchaindb.com/docs/sdk)[Natural-language queries](https://originchaindb.com/docs/ask) Keep exploring ## Answers, updates and conversations. Find the right next step for your team. [### Common questions Product, configurations and getting started.](https://originchaindb.com/faq)[### Events & technical sessions Learn with the team and the community.](https://originchaindb.com/events)[### Talk through your workload Bring your data model and the queries you need.](https://originchaindb.com/contact) --- # Database security and isolation | OriginChainDB Canonical source: https://originchaindb.com/security Sitemap last modified: 2026-09-23T08:33:23.000Z Security # Clear controls. Defined responsibilities. Understand the request path, identity model and deployment controls before putting your application into production. [Read the access reference](https://originchaindb.com/docs/dashboard/iam#enforced)[Report a vulnerability](https://originchaindb.com/security#responsible-disclosure) Public transportHTTPS endpoints IdentityApplication and database access IsolationShared or dedicated deployment ## Follow the request. Keep the identity clear. Application authorization and database permissions are separate parts of the design. Application and database responsibilitiesIllustrative HTTP request path Your applicationOriginChainDB 01 / User interface ### Browser or mobile app Authenticate your users and send requests to your application server. User session 02 / Application server ### Keep credentials here Apply application access rules and choose the intended database identity. Server-side secret storage 03 / Database endpoint ### Authenticate the request The credential determines the database identity and applicable query checks. Query-specific enforcement Public request`HTTPS + Authorization: Bearer `Credential scope depends on its type and deployment. Which identity reaches the database? Administrative bearer The master bearer acts as a database administrator. It is not a test of a restricted user's permissions. Database user Database-user credentials carry that user's roles. Validate the documented enforcement for each query path. Organization member Console membership is a separate identity layer; it does not itself grant database-user permissions. [Read the enforcement reference](https://originchaindb.com/docs/dashboard/iam#enforced) Example HTTP application pattern. Keep full-access database bearers out of browser and mobile client code. Deployment isolation ## The configuration matters. Dedicated infrastructure and shared hosting have different isolation and recovery properties. Free and Starter ### Shared infrastructure These databases run on shared, scale-to-zero infrastructure. Each database has its own hostname and bearer credential. For Free, public TLS ends at the shared load balancer; traffic then travels over the private network to the database. Dedicated configurations ### Dedicated compute and storage Virtual machines and database storage serve one customer. An optional standby is a separate deployment choice. Dedicated configurations: encrypted nightly snapshots, included. Standby replication is asynchronous. Review the configured backup path and current recovery caveats separately. [Deployment reference](https://originchaindb.com/docs/deploy#architecture)[Durability and recovery](https://originchaindb.com/docs/ops#durability) Access and operational records ## Choose the control for the request. Use the deployed service's capabilities and the intended database identity when evaluating access. [Credential handling guide](https://originchaindb.com/docs/auth) ### Credential handling Keep full-access bearers in server-side secret storage. Plan credential replacement across clients, and revoke exposed credentials. [Rotation procedure](https://originchaindb.com/docs/auth#rotation) ### Database-user reads The documented read controls include table grants, row policies and column masking. Test the intended identity and each query path. [Policies and identities](https://originchaindb.com/docs/dashboard/iam#rls) ### Write restrictions Write-grant letters alone do not establish a write barrier. Use service-enforced read-only credentials or a separate database where a write boundary is required. [Enforcement scope](https://originchaindb.com/docs/dashboard/iam#enforced) ### Authorization audit The engine records user, role, grant and policy changes in a hash-linked audit chain. Console preview history is a separate record. [Audit reference](https://originchaindb.com/docs/dashboard/iam#cluster-audit) Responsible disclosure ## Report a security issue. Send a minimal reproduction, the affected version, and your assessment of severity. [security@originchain.ai](mailto:security@originchain.ai) ### Encrypted disclosure For encrypted disclosure, [request a PGP key by email](mailto:security@originchain.ai?subject=PGP%20key%20request) before sending sensitive details. ### Reporter credit Reporters are credited by name in release notes unless they prefer to remain anonymous. ### Published response targets 24 hours First human response acknowledging your report. 7 days Triage: severity, reproducibility, and a fix plan. 30 days Patch released or a public advisory, whichever comes first. Disclosure scope and exclusions Scope is the OriginChainDB managed cloud, the engine binary, and the customer-facing console. The static marketing site is out of scope. ### In scope - Remote code execution in any ingress path (HTTPS, SSE, the bearer-auth layer). - Auth or access-control bypass in the engine or the managed platform. - Tenant-to-tenant crossover on the managed cloud (network, identity, or process boundary). - Data exfiltration or corruption via a crafted query or row write. - LLM prompt-injection that escapes the plan compiler and reaches storage. - Cryptographic or integrity flaws in the durability, recovery, full-text, vector, or backup subsystems. - Bypass of per-tenant rate-limit / quota or per-API-key bucket accounting. - Concurrency hazards that violate single-row CAS or schema-cutover atomicity. ### Out of scope - Denial-of-service from obviously abusive query volume within your configuration's quota. - Issues requiring a compromised host or physical access to disk. - Vulnerabilities in third-party dependencies already tracked by their maintainers. - Social engineering of OriginChainDB staff or support. - The marketing site (originchaindb.com) — it is a static deployment. ## Review your deployment with us. For questionnaires, current assessment status or contractual requirements, contact the security team. [Talk to security](mailto:security@originchain.ai)[Trust resources](https://originchaindb.com/trust)[Technical architecture](https://originchaindb.com/architecture) --- # Service levels - availability, recovery, support Canonical source: https://originchaindb.com/sla Sitemap last modified: 2026-09-23T08:33:23.000Z service levels # What we commit to, per configuration. Monthly uptime is best-effort on a single zone, 99.9% on HA, 99.95% on HA+ and 99.99% on Enterprise; the remedy is a pro-rata service credit against the next invoice. One table below for availability, one for recovery, one for support — every figure is the same figure on /pricing, in the FAQ and in the Terms. This page summarises, [the Terms of Service](https://originchaindb.com/terms) govern. 99.9% 1 monthly uptime · HA 99.95% 2 monthly uptime · HA+ 99.99% 3 monthly uptime · Enterprise 5 s 4 recovery boundary · PITR 1. 1 Contractual — HA configuration. /terms § Availability — monthly uptime on HA (two zones, writer + standby). 2. 2 Contractual — HA+ configuration. /terms § Availability — monthly uptime on HA+. 3. 3 Contractual — Enterprise agreement. /terms § Availability — monthly uptime on Enterprise. 4. 4 Engine changelog 2026-08: restore lands on a roll-up boundary, 5 s by default; the Intra-Segment PITR add-on tightens it to 500 ms. 01 · availability ## Monthly uptime targets and what you get if we miss. | Configuration | Topology | Monthly uptime | Remedy | | --- | --- | --- | --- | | Single-zone | One dedicated instance in one zone | best-effort | — | | HA | Two zones · writer + cross-zone standby · automatic failover | 99.9% | pro-rata service credit against the next invoice | | HA+ | Multi-zone · writer + two standbys · higher-capacity configuration | 99.95% | pro-rata service credit against the next invoice | | Enterprise | As HA+, under an annual agreement with negotiated quotas | 99.99% | pro-rata service credit against the next invoice | Uptime is measured per instance per calendar month. Scheduled maintenance, force-majeure events and outages caused by your own configuration are excluded from the calculation, as defined in the Terms. Credits are the sole remedy for availability shortfalls. 02 · recovery ## Recovery objectives you can put in a runbook. Backups encrypted nightly snapshots, included, integrity-checked daily, kept in the instance's region. Point-in-time recovery opt-in. Restore lands on a 5 s roll-up boundary; 500 ms with the Intra-Segment PITR add-on. Retention window 7 or 30 days self-serve, up to 35 days. Longer windows are agreed per deployment on Enterprise. Restore path From the console or the HTTP API. 03 · support ## Support that reaches an engineer. Support reaches us at [support@originchain.ai](mailto:support@originchain.ai). Enterprise agreements include a named support engineer, and their response-time commitments are written into the agreement. ## Need a different number? Ask for it in writing. Custom availability, retention and support terms are agreed on Enterprise contracts. [Discuss an enterprise agreement](https://originchaindb.com/contact?topic=enterprise) [Read the Terms](https://originchaindb.com/terms) --- # Application patterns for a multimodal database | OriginChainDB Canonical source: https://originchaindb.com/solutions Sitemap last modified: 2026-09-24T06:54:16.000Z Solutions / Distributed multimodal database # One database. Many ways to build. Turn records, meaning and relationships into useful application workflows. [Find your use case](https://originchaindb.com/solutions#solution-paths)[Meet the database](https://originchaindb.com/product/database) ## Start with the work your app needs to do. Select a workflow to see where the database fits. Six workflows. The same four data models.Interactive example ### Answers with sources. [Explore RAG](https://originchaindb.com/solutions/rag) 01 / YOUR DATAArticles & passages`document_id · content · source` 02 / ORIGINCHAINDB SQLVectorGraphFull-text Query the models this workflow needs. 03 / YOUR APPLICATIONRetrieve → merge → assemble context Passages with source references Your orchestration. Your decisions. Illustrative architecture · no queries are executed.[How the database connects the models](https://originchaindb.com/product/database) Explore the patterns ## A practical path for each application. ### [RAG](https://originchaindb.com/solutions/rag) Retrieve relevant passages and connect them to the records your application can cite. VectorFull-textSQL ### [AI agents](https://originchaindb.com/solutions/agents) Store structured state, retrieve useful memories and trace the records behind a run. SQLVectorGraph ### [Recommendations](https://originchaindb.com/solutions/recommendations) Find similar items, follow relationships and apply your catalog's business rules. VectorGraphSQL ### [Fraud detection](https://originchaindb.com/solutions/fraud-detection) Connect accounts, devices and transactions so analysts can inspect the evidence. GraphSQL ### [Personalization](https://originchaindb.com/solutions/personalization) Combine explicit preferences and relevant content with application-owned eligibility rules. SQLVectorFull-text ### [Real-time analytics](https://originchaindb.com/solutions/real-time-analytics) Filter, group and inspect records in the database your application already writes to. SQL One data layer ## Choose the models. Own the workflow. [Explore the architecture](https://originchaindb.com/architecture) OriginChainDB stores and queries Records, embeddings, search indexes and declared relationships, connected through stable record IDs. Your application coordinates Separate writes and queries, ranking, model calls and the rules that decide what to show or do next. Start with a small evaluation Use representative data and explicit checks for relevance, permissions and query behavior before scaling. ## Bring your first workflow to life. [Start free](https://app.originchaindb.com/signup)[Follow the quickstart](https://originchaindb.com/docs/quickstart) --- # Persistent context for AI agents | OriginChainDB Canonical source: https://originchaindb.com/solutions/agents Sitemap last modified: 2026-09-24T06:54:16.000Z [← All solutions](https://originchaindb.com/solutions) AI agents # Give the next turn something to remember. Keep run state, useful memories and tool outcomes in one database. Retrieve the context your agent needs, with records you can inspect. [Start free](https://app.originchaindb.com/signup)[Build agent memory](https://originchaindb.com/docs/examples/atomic-multi-shape/agent-memory) ## Follow one turn, from recall to record. Your agent runtime controls the loop. OriginChainDB stores the state and context it reads and writes along the way. Inspect a support agent's turnInteractive illustration Application-designed workflowA support agent’s next turn `session-84 / run-12` Your runtime starts the next turn. OriginChainDB · example record`agent.memories` Record ID memory-03 Session session-84 Observation Customer prefers weekly updates. Recall matching memory IDs, then load their source records before building the next prompt. Linked by application IDs - Run state - Memory records - Tool events Rows, embeddings and text indexes use separate writes. Illustrative records. Select a stage to inspect its role; no agent or tool is executed. A suggested data model ## Separate the state. Keep the relationships. Use session, run and record IDs to connect observations with their source. These are application tables you define, not a required agent framework. 01 ### Run state Current step, status and checkpoint data. `run_id · session_id · status` 02 ### Memory A focused observation with its time and source. `memory_id · agent_id · text` 03 ### Tool event The attempted action and recorded outcome. `call_id · run_id · outcome` Add embeddings and full-text entries under the memory's ID. Row, vector and text writes are separate calls; track which steps have completed. [See the memory schema and writes](https://originchaindb.com/docs/examples/atomic-multi-shape/agent-memory) ## Recall by the question you need to answer. Use structured lookups for state, similarity for related memories, keywords for exact terms and graph references for provenance. [SQLWhat is this run waiting for?](https://originchaindb.com/product/sql)[VectorHave we seen a similar request?](https://originchaindb.com/product/vector)[Full-textWhere did error AX-7 appear?](https://originchaindb.com/product/fts)[GraphWhich source produced this memory?](https://originchaindb.com/product/graph) Tools with a clear contract ## Let the agent ask. Inspect what it runs. Expose focused database tools through your application. ASK accepts a question, table scope and optional plan inspection, then returns database rows. It is a query interface. Your runtime still chooses tools, calls models and decides whether an action needs approval. [Read the ASK contract](https://originchaindb.com/docs/ask) Example request`POST /v1/tenants/:tenant/ask` ``` { "nl": "agent.memories where session_id = 'session-84' limit 5", "schemas": ["agent.memories"], "show_plan": true } ``` Scope the request to the memory table and review its returned plan. ## Make retries and retention explicit. Durable records are the starting point. Define how your application handles partial writes, stale context and access to each session. Retry deliberately Keep stable memory and tool-call IDs. Retry failed index writes without treating a repeated tool request as a new external action. Expire the whole memory When retention ends, remove the row and its vector and text entries. Row deletion does not automatically remove those indexes. Scope every lookup Use database access controls and validate agent/session scope in your service. Carry agent metadata into retrieval indexes; filters do not replace authorization. ## Build on memory you can inspect. [Start the memory example](https://originchaindb.com/docs/examples/atomic-multi-shape/agent-memory)[Explore the database](https://originchaindb.com/product/database) --- # Investigate connected fraud patterns | OriginChainDB Canonical source: https://originchaindb.com/solutions/fraud-detection Sitemap last modified: 2026-09-24T06:54:16.000Z [All solutions](https://originchaindb.com/solutions) Fraud investigation · SQL + graph # Follow the records. Investigate the connection. Query transactions and trace related accounts, devices and beneficiaries in the same database. [Start free](https://app.originchaindb.com/signup)[Build your first traversal](https://originchaindb.com/docs/graph/quickstart) ## A connection is a starting point. A shared device or beneficiary adds context to an account. Inspect the underlying records and their timing before deciding what the connection means. Account relationships and supporting evidenceInteractive illustration Follow a connection account-AStarting account device-31Device reference account-BRelated account payee-08Beneficiary Inspect the link Evidence · device reference ### Two accounts. One recorded device. Both account records reference `device-31`. That connection is a reason to investigate, not a finding of fraud. Starting point account-A Related record account-B Next question Was this a shared household device? Highlighted lines show the selected relationship. All records are illustrative. From signal to case ## Let the query explain why you looked. OriginChainDB retrieves records and relationships. Your application defines the review criteria, assembles the case and routes it to an analyst. 1. 01 ### Select a candidate Use SQL to group transfers by account and examine counts or amounts. Your application chooses thresholds and the time window. 2. 02 ### Expand the context Follow declared relations from the candidate. Bound the number of hops and retain the identifiers behind each connection. 3. 03 ### Make a reviewable decision Attach source records, query parameters and timestamps to the case. Approval, escalation and account restrictions remain application decisions. ## Ask a specific question. Choose the query for the evidence you need. A graph pattern describes a relationship; it does not establish intent. Transaction activity ### Which accounts need a closer look? Group transfers with SQL. Compare counts and totals using your application's review rules. [SQL aggregates](https://originchaindb.com/docs/sql#groupby) Shared connections ### Who else references this device? Inspect neighboring records and reverse relations to trace a shared identifier. [Neighbors and relations](https://originchaindb.com/docs/graph#neighbors) Reachability ### How are these accounts connected? Use bounded traversal or path queries, with explicit depth and result limits. [Bounded traversal](https://originchaindb.com/docs/graph#bfs) Keep the evidence traceable ## Model facts. Preserve their source. Give accounts, devices and transfers stable IDs. Declare relations on columns holding the target record's key. [Define a graph relation](https://originchaindb.com/docs/schemas/graph) Accounts Account ID and device reference Transfers Transfer ID, source, beneficiary, amount and timestamp Case evidence Application-owned query context and review outcome ## Bound the search. Keep the judgment. Graph capability must be enabled for the selected instance. Start with a narrow frontier; broad paths can exceed query budgets. Review the documented limits before expanding a case. Shared identifiers can have ordinary explanations. Verify identity, timing and source quality, and let your application or analyst decide whether to act. [Review graph limits](https://originchaindb.com/docs/graph#graph-limits) ## Build an investigation you can explain. [Start with graph relations](https://originchaindb.com/docs/graph/quickstart)[Explore the graph module](https://originchaindb.com/product/graph) --- # Personalization from application data | OriginChainDB Canonical source: https://originchaindb.com/solutions/personalization Sitemap last modified: 2026-09-23T08:33:23.000Z [← All solutions](https://originchaindb.com/solutions) Personalization # An experience that reflects what people choose. Connect session preferences, content and retrieval signals. Your application decides which experience to build from them. [Start free](https://app.originchaindb.com/signup)[Start the integration](https://originchaindb.com/docs/quickstart) Context becomes a useful experience ## Same catalog. A different starting point. Change one explicit interest in this sample session. The application selects a different set of eligible articles without changing the underlying content catalog. Preference-driven content selectionInteractive illustration Illustrative sessionChange the expressed interest. Session context Interest Outdoor guides Language English Experience Beginner Personalization choice Enabled for this example Your application supplies these preferences. An application-curated feed3 eligible articles 1. Plan your first day hikeA starting point for beginners 2. Pack light for the trailMatches the selected interest 3. Read a trail mapA related skill to explore Application eligibility checks - Published - English - Allowed audience Context → candidates → eligibility → ordering The application chooses the ranking and assembles the feed. Local sample. No profile is saved. OriginChainDB stores and retrieves the data. Consent handling, eligibility rules and ranking belong to your application. [Explore retrieval patterns](https://originchaindb.com/docs/rag) Choose useful context ## Start with what the person tells you. Build from the signals your experience needs. Keep their meaning explicit so a preference, an interaction and an access rule do not become interchangeable. Declared preferences Store selected topics, language and experience level as structured fields. A visitor can change these directly. Session interactions Record relevant events under your application's policy. Your code decides whether an event should change future retrieval. Content representations Keep searchable text, embeddings and declared relationships linked to content IDs, so candidate results resolve to their records. Eligibility leads. Ranking follows. ## Start with what can be shown. Personalization does not replace authorization. Apply the relevant database identity and access controls, then enforce application rules before displaying content. Eligible content ### Published and appropriate - Check publication status. - Apply language and audience requirements. - Respect the visitor's personalization choice. [Read access controls](https://originchaindb.com/docs/auth) Application ordering ### Relevant and useful Combine declared interests, retrieval ranks and any editorial priorities you choose. Keep a reason for each selection so teams can inspect the experience. Fetch candidate records with SQL for structured checks. Vector filtering is index-specific; it is not a substitute for the full eligibility policy. [Check vector filtering](https://originchaindb.com/docs/vector) A preference can change ## Update deliberately. Fall back gracefully. A row update does not also replace its embedding or full-text entry. Coordinate those calls explicitly and track which representation has been refreshed. [Review the write workflow](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog) 1. 01 ### Save the explicit choice Write the updated preference fields and record the application's profile version. 2. 02 ### Refresh derived representations If your design uses a preference embedding, generate and write the replacement separately. 3. 03 ### Choose a fallback When context is missing, disabled or awaiting refresh, serve an eligible curated selection. Evaluate the whole experience ## Make the outcome observable. Compare application recipes with a consistent evaluation set. Record enough context to explain a result without retaining unrelated session data. Coverage ### Can you fill the experience? Count eligible results and fallback use across different interests and audience rules. Relevance ### Did the selection help? Review examples with your team and measure the outcome your application is designed to improve. Freshness ### Which context was used? Track profile versions and representation updates alongside the application's ranking recipe. [Model session data](https://originchaindb.com/docs/schemas)[Retrieve similar content](https://originchaindb.com/docs/vector/quickstart)[Explore product recommendations](https://originchaindb.com/solutions/recommendations) Build with OriginChainDB ## Give context a useful role. [Start free](https://app.originchaindb.com/signup)[Discuss your application](https://originchaindb.com/contact) --- # Hybrid retrieval for RAG applications | OriginChainDB Canonical source: https://originchaindb.com/solutions/rag Sitemap last modified: 2026-09-24T06:54:16.000Z [← All solutions](https://originchaindb.com/solutions) Retrieval-augmented generation # Find the context. Keep the source. Build answers from relevant passages, exact matches and source records in one database. Your application decides what reaches the model. [Start free](https://app.originchaindb.com/signup)[Build a RAG workflow](https://originchaindb.com/docs/rag) ## Two ways to find the right passage. Semantic search finds related meaning. Keyword search preserves names, product codes and exact phrases. Explore how their results become context. From a question to cited contextInteractive illustration Your application’s questionHow do I claim a warranty for AX-7? Your app embeds the question and also sends its search text. OriginChainDB · two retrieval calls Vector search ### Find similar meaning 1. policy:12Repair and replacement cover 2. guide:04Starting a service request Embedding → related passages Full-text search · BM25 ### Find the exact terms 1. manual:07AX-7 warranty claims 2. policy:12Repair and replacement cover Search text → matching passages Ranked IDs return to your application Your application ### Merge. Verify. Build context. Combine both lists by rank, remove duplicates and fetch eligible source rows. Generation stays in your application. Example context, with source references 1. [1] AX-7 warranty claimsmanual:07 · page 7 2. [2] Repair and replacement coverpolicy:12 · page 12 Pass selected passages to your model with their citations. Illustrative rankings and records. These controls explore the workflow locally. Prepare the knowledge ## One passage. A shared record ID. Split documents into useful passages. Store the text, source location and document version; use the same passage ID for retrieval indexes. [Follow the ingestion example](https://originchaindb.com/docs/rag#ingest) Example passage ID`manual:07` 1. 01 ### Save the source row Text, document ID, page and version. 2. 02 ### Put its embedding Generate the vector with your chosen embedding model. 3. 03 ### Index searchable text Add the passage to the full-text index. These are separate API calls, each atomic on its own. Track completion and retry failed steps before publishing updated content. ## Choose the retrieval combination you need. Text and meaning ### Vector search + BM25 Run vector and full-text queries separately. Your application merges the ranked IDs, removes duplicates and loads source rows. Reciprocal Rank Fusion is one way to combine the lists. [Build the retrieval flow](https://originchaindb.com/docs/rag#retrieve) Two embedding representations ### Dense + sparse vectors The vector hybrid endpoint fuses dense and sparse vector rankings on the server. Its sparse inputs are weighted vector components; this route does not query the BM25 text index. [Explore vector hybrid search](https://originchaindb.com/docs/examples/vector/hybrid-topk) Before the answer ## Give every passage a way back. Keep citations attached as results move through retrieval, reranking and prompt construction. A fluent answer still needs evidence. Source Retain the document ID, page or section, and version used for the answer. Eligibility Apply access checks before passing source text to the model. Metadata filters help narrow retrieval; they do not replace authorization. Coverage Evaluate with known questions and expected sources. When context is insufficient, let the application ask for clarification. ## Your workflow. Your generation model. OriginChainDB stores and queries the retrieval data. Your application handles chunking, embeddings, result merging, prompts, model calls and answer evaluation. ### Where ASK fits ASK translates natural-language questions into database queries and returns rows. It can serve as a query tool within your application; it does not run a complete RAG generation pipeline. [See the ASK interface](https://originchaindb.com/product/ask) ## Build an answer you can trace. [Start the RAG walkthrough](https://originchaindb.com/docs/rag)[Explore vector search](https://originchaindb.com/product/vector) --- # Analytics on operational records | OriginChainDB Canonical source: https://originchaindb.com/solutions/real-time-analytics Sitemap last modified: 2026-09-24T06:54:16.000Z [All solutions](https://originchaindb.com/solutions) Operational analytics · SQL # Your application data. A clearer view of the work. Count orders, track queues and summarize activity from the records your application writes. [Start free](https://app.originchaindb.com/signup)[Run your first SQL query](https://originchaindb.com/docs/sql/quickstart) ## Six records. One useful summary. A grouped SQL query turns individual orders into counts and totals. Your application renders the returned rows as a chart, table or alert. Source rows → SQL aggregation → application chartIllustrative dataset · USD 01 ### Operational records shop.orders | id | status | amount_usd | | --- | --- | --- | | o-101 | paid | 120 | | o-102 | paid | 80 | | o-103 | pending | 60 | | o-104 | paid | 200 | | o-105 | pending | 140 | | o-106 | failed | 50 | 02 ### Group and aggregate ``` SELECT status, COUNT(*) AS orders, SUM(amount_usd) AS total_usd FROM shop.orders GROUP BY status; ``` One result row per status, with its order count and total order value. 03 ### Render the result Number of orders paid 3 pending 2 failed 1 6 source rows → 3 status groups. Select order value to compare the sums. From business question to SQL ## Define the metric before the dashboard. Choose the rows that belong in the calculation, the grouping key and the meaning of the result. [Read the aggregate guide](https://originchaindb.com/docs/sql#groupby) What is waiting? Filter pending orders, then count them by region or queue. Keep completed records outside that metric. What has been paid? Filter paid orders before summing their value. A total across all statuses is not paid revenue. Which groups need attention? Use HAVING to select groups after aggregation. Your application decides when a result creates an alert. A plan you can inspect ## See what the query asks the engine to do. EXPLAIN returns a plan without executing the SELECT. Inspect scans, filters and aggregates before putting a query behind a frequently refreshed dashboard. If you expected an index scan, review the schema's index declarations and the predicate. Measure against your own data and instance configuration. [Read query plans](https://originchaindb.com/docs/sql#explain) Inspect the pending-order query ``` EXPLAIN SELECT status, COUNT(*) AS orders FROM shop.orders WHERE status = 'pending' GROUP BY status; ``` Run this on your instance to inspect the operators chosen for your schema and indexes. ## Freshness includes the application. Querying operational records removes a separate analytical copy from this flow. The dashboard still needs a clear refresh policy. 01 · Database ### Query the relevant records Use explicit filters and the correct endpoint. Check transaction and replica behavior for the deployment you use. 02 · Application ### Choose a refresh interval Schedule requests, cache intentionally and display when the last successful query completed. 03 · Viewer ### Handle the next update Keep the previous result during refresh and surface failures. A stale chart should not appear current. Keep the workload focused ## Operational questions, with an execution budget. Use this approach for application dashboards, order summaries and workload queues. For large historical scans, evaluate a separate analytical workflow against your requirements. Sorts, joins and window functions can buffer their inputs. Narrow the dataset before expensive operations; LIMIT after ORDER BY does not remove that work. Handle query limits and retry responses in the application. [Review SQL execution limits](https://originchaindb.com/docs/sql#limits) ## Start with the question your team asks every day. [Try an operational query](https://originchaindb.com/docs/sql/quickstart)[Explore the SQL module](https://originchaindb.com/product/sql) --- # Recommendations using similarity and relationships | OriginChainDB Canonical source: https://originchaindb.com/solutions/recommendations Sitemap last modified: 2026-09-23T08:33:23.000Z [← All solutions](https://originchaindb.com/solutions) Recommendations # Relevant products. Clear reasons. Combine similarity, product relationships and business rules to build a shortlist your application can explain. [Start free](https://app.originchaindb.com/signup)[Try vector retrieval](https://originchaindb.com/docs/vector/quickstart) The recommendation recipe 1. DiscoverSimilar and related candidates 2. QualifyCurrent stock, price and availability 3. OrderA ranking your application controls From a product to a shortlist ## Different signals. One application experience. Retrieve candidates through the vector and graph APIs, fetch their product rows, then apply your ranking. The example changes the application recipe while keeping eligibility fixed. A sample outdoor catalogInteractive illustration Starting productTrail runner product-9 Vector similarityDeclared accessory links ### Retrieve candidates - Ridge runnerSimilar embedding · product-12 In stock - All-weather runnerSimilar embedding · product-18 Sold out - Light daypackAccessory relation · product-22 In stock - Running socksAccessory relation · product-31 In stock Your application ### Check. Combine. Rank. 1. Fetch current rows 2. Remove unavailable items 3. Order eligible candidates ### Choose a ranking recipe 1. 01 Ridge runnerSimilar embedding 2. 02 Light daypackAccessory relation 3. 03 Running socksAccessory relation Sample ordering set by the application. No live query or measured scores. Your application coordinates these requests. This is not a single cross-model query or a shared-snapshot guarantee. [See retrieval patterns](https://originchaindb.com/docs/rag#retrieve) Prepare the catalog ## Give each signal a record to point to. 01 · Products ### The business facts Store names, prices, stock and availability in rows. A stable product ID connects those fields to every retrieval result. [Define your schema](https://originchaindb.com/docs/schemas) 02 · Embeddings ### What feels similar Generate embeddings with your chosen model. Keep the dimensions and distance metric consistent across product writes and queries. [Configure vectors](https://originchaindb.com/docs/schemas/vector) 03 · Relations ### What belongs together Declare accessory or category relations using record references. Your application decides which connections contribute useful candidates. [Model relationships](https://originchaindb.com/docs/schemas/graph) Eligibility before ranking ## A close match still needs to qualify. Use SQL to check retrieved IDs against the fields that determine whether a product can be shown. Keep merchandising rules separate from similarity scores. - Remove unavailable products. - Apply market and price requirements. - Deduplicate IDs before ordering. Vector metadata filters use exact equality and behave differently by index. IVF-PQ does not apply them; check the index contract before relying on a filter. [Read filtering behaviour](https://originchaindb.com/docs/vector) Fetch eligible candidate rowsSQL example ``` SELECT id, name FROM shop.products WHERE id IN ($1, $2, $3) AND in_stock = true AND price_cents <= $4 LIMIT 12; ``` Bind candidate IDs and the budget through the request's `params` array. Keep the experience current ## Plan the update path, too. Row, vector and full-text updates are separate calls. Update the affected representations when a product changes, and track any stage that needs retrying. 1. 01 ### Write the record Change business fields and declared relationship references. 2. 02 ### Refresh derived data Replace embeddings or search entries when their source content changes. 3. 03 ### Handle partial progress Retry failed stages and recheck purchase availability in the application. [Follow catalog ingestion](https://originchaindb.com/docs/examples/atomic-multi-shape/product-catalog) Build and evaluate ## Judge the shortlist, not just the search. Compare recipes on representative products. Track eligible result coverage, relevance and latency across the entire application flow. Keep a curated fallback when too few candidates survive. [Retrieve the first candidatesVector quickstart](https://originchaindb.com/docs/vector/quickstart)[Add relationship contextGraph methods and limits](https://originchaindb.com/docs/graph)[Review the measurementsPublished workloads and methodology](https://originchaindb.com/benchmarks) Build with OriginChainDB ## Make the next suggestion useful. [Start free](https://app.originchaindb.com/signup)[Talk through your workload](https://originchaindb.com/contact) --- # Terms of service - OriginChain managed cloud Canonical source: https://originchaindb.com/terms Sitemap last modified: 2026-09-27T14:07:04.000Z terms # Terms of service These terms describe how the OriginChain managed cloud, this website, and our name may be used. Operated by Silicoyn Technologies Pvt Ltd from Bengaluru. Written in English, not Latin. last updated 27 September 2026 Browse this document - [Scope of these terms](https://originchaindb.com/terms#scope) - [Eligibility](https://originchaindb.com/terms#eligibility) - [The managed cloud service](https://originchaindb.com/terms#service) - [Acceptable use](https://originchaindb.com/terms#acceptable-use) - [Service levels](https://originchaindb.com/terms#sla) - [Your data](https://originchaindb.com/terms#data) - [No warranty (beyond the SLA)](https://originchaindb.com/terms#no-warranty) - [This website](https://originchaindb.com/terms#website) - [The name and mark](https://originchaindb.com/terms#trademark) - [Limitation of liability](https://originchaindb.com/terms#liability) - [Indemnity](https://originchaindb.com/terms#indemnity) - [Grievance officer](https://originchaindb.com/terms#grievance) - [Governing law and jurisdiction](https://originchaindb.com/terms#law) - [Changes to these terms](https://originchaindb.com/terms#changes) ## Scope of these terms These terms cover two things: the marketing website at originchain.ai, and the managed cloud service ("the Service") operated by Silicoyn Technologies Pvt Ltd, an Indian private limited company with its registered office in Bengaluru, Karnataka. Different clauses apply to each - they are called out where they diverge. Silicoyn Technologies Pvt Ltd is the operating company behind the OriginChain product brand. References below to "we", "us", and "our" mean Silicoyn Technologies Pvt Ltd. ## Eligibility To enter into these terms you must be at least 18 years old, have legal capacity to contract under the Indian Contract Act 1872, and not be barred from receiving services under the laws of India. If you sign up on behalf of an organization, you represent that you have authority to bind that organization, and "you" then means both you and the organization. ## The managed cloud service OriginChain Cloud is a subscription service. By signing up you authorize us to provision resources in the region you pick and to bill you monthly at the list price of the configuration you choose. Upgrades take effect immediately; downgrades take effect at the next billing cycle. A valid payment method is captured at instance creation and billing begins from day 1 on the configuration you select. You may cancel at any time by deleting the configuration in the console. Deleting it stops the service and deletes the database straight away; this cannot be undone. Its stored data is erased within 7 days and its backups within 30 days. See our [Cancellation & Refund Policy](https://originchaindb.com/refund-policy). Your card stays on file as long as you have one or more provisioned instances. To remove the card you must first delete every instance and let the 30-day data retention window expire. This is a deliberate constraint of the subscription model: an active subscription requires a card on file to retry failed charges and to remit overage usage. Payments are processed by Razorpay Software Pvt Ltd (an RBI-licensed payment aggregator) under their own terms. We never receive your card number, CVV, or expiry - only a tokenised reference. You are responsible for how you use the Service - the rows you put in, the questions you ask, the people you give the bearer token to. Do not use it to store unlawful content or to attack third parties. We reserve the right to suspend accounts for flagrant abuse, fraud, or activity that violates Indian law. Either party may terminate the subscription at any time. On termination, every row and backup of your database is deleted within 30 days. Billing records are kept for the period required by Indian tax law (currently 7 years under the Income Tax Act 1961). ## Acceptable use You will not use the Service to: (a) store, transmit, or process content that is unlawful under Indian law (Information Technology Act 2000 Sections 67, 67A, 67B and the IT Rules 2021); (b) infringe any intellectual-property right; (c) impersonate another person; (d) probe, scan, or test the vulnerability of any system without our prior written consent; (e) attempt to bypass any rate limit, quota, or isolation boundary; or (f) use automated means to scrape data not belonging to you. You agree to promptly remove any content that we, in good faith, identify as violating this section after notifying you. This is consistent with our role as an intermediary under the IT Rules 2021. ## Service levels The Service targets best-effort uptime on Single-zone, 99.9% monthly uptime on HA, 99.95% on HA+, and 99.99% on Enterprise. If we miss the target on a paid configuration, you receive a pro-rata service credit against your next invoice per the SLA published at /docs#sla. Credits are the sole and exclusive remedy for availability shortfalls. Scheduled maintenance, force-majeure events, and outages caused by your own configuration or by your own account-level state (billing holds, access mis-configurations) are excluded from the SLA calculation. ## Your data You own every byte you put into an OriginChain instance. We do not claim any license over it. We process it solely to provide the Service to you, in the region you picked, under the Data Processing Addendum at /dpa. We will not sell, share, or train models on your data. Period. Personal data is processed in accordance with the Digital Personal Data Protection Act 2023 ("DPDP Act") and our Privacy Notice at /privacy. You retain all rights of a Data Principal under that statute and may exercise them by writing to Silicoyn Technologies Pvt Ltd at the contact in the Privacy Notice. ## No warranty (beyond the SLA) Except for the SLA above, the Service is provided "as is," without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, and non-infringement. To the extent the Consumer Protection Act 2019 applies, nothing in this clause limits a statutory right that cannot be excluded by contract. You are responsible for evaluating whether OriginChain fits your workload and for benchmarking it on your data. We will help you do that - but the decision and the outcome are yours. ## This website The content of originchain.ai is provided for informational purposes. Benchmarks, latency figures, and architecture claims describe our reference deployment; your results will depend on your hardware, data, and workload. We reserve the right to update any page on this site without notice. Material changes will be reflected in our changelog. ## The name and mark "OriginChain" and the OriginChain wordmark are trademarks of Silicoyn Technologies Pvt Ltd. You may use them to describe or refer to the Service ("running on OriginChain", "built on OriginChain"). You may not use them to imply endorsement of a competing hosted service or a derivative product without our written permission. Other names used on this site belong to their owners. Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries. PostgreSQL is a trademark of the PostgreSQL Community Association of Canada. MySQL is a registered trademark of Oracle Corporation and/or its affiliates. Oracle is a registered trademark of Oracle Corporation and/or its affiliates. Microsoft and SQL Server are registered trademarks of Microsoft Corporation in the United States and/or other countries. Apache Lucene is a trademark of the Apache Software Foundation. We refer to them only to describe compatibility with those interfaces, which our software implements independently. We are not affiliated with, endorsed by, or sponsored by any of them, and we do not distribute their software. ## Limitation of liability To the maximum extent permitted by law, Silicoyn Technologies Pvt Ltd will not be liable for any indirect, incidental, special, consequential, or punitive damages arising out of your use of the Service or this website - even if advised of the possibility of such damages. For direct damages arising out of the Service, our aggregate liability is capped at the fees you paid to us in the twelve months preceding the claim. Nothing in this clause limits liability for fraud, gross negligence, or any other liability that cannot be excluded under Indian law. ## Indemnity You agree to indemnify and hold us harmless against third-party claims arising from your content stored on or processed by the Service, your breach of the Acceptable Use clause, or your violation of any law applicable to your use of the Service. ## Grievance officer In compliance with the Information Technology (Intermediary Guidelines and Digital Media Ethics Code) Rules 2021 and Section 7 of the DPDP Act 2023, the Grievance Officer / Data Protection Officer for Silicoyn Technologies Pvt Ltd is reachable at grievance@originchain.ai. Complaints are acknowledged within 24 hours and resolved within 15 days. ## Governing law and jurisdiction These terms are governed by the laws of India. The parties submit to the exclusive jurisdiction of the courts at Bengaluru, Karnataka, India for any dispute arising out of or in connection with these terms. Notwithstanding the above, either party may seek interim or injunctive relief in any court of competent jurisdiction to protect its intellectual property or confidential information. ## Changes to these terms We may update these terms from time to time. Material changes - anything that meaningfully reduces your rights or increases your obligations - will be announced at least 30 days in advance via email to the address on file and via our changelog at /changelog. Non-material changes (clarifications, typo fixes, references to new sub-pages) take effect on posting. Continued use of the Service after a change constitutes acceptance of the updated terms. ## Need something clarified? Volume contracts, data residency, a BAA, an unusual deployment scenario - we'll give you a straight answer, not a lawyer one. [support@originchain.ai](mailto:support@originchain.ai) --- # Published database test reports | OriginChainDB Canonical source: https://originchaindb.com/tests Sitemap last modified: 2026-09-23T08:33:23.000Z live evidence # Published database test reports Inspect published test runs, pass and fail counts, latency measurements and the underlying reports. 01 · headline numbers ## The most recent run. No spin. Pulled from `/api/test-reports/index.json` on page load. The numbers below are the latest published report's metrics. total tests - cargo test workspace failures - zero is the only acceptable answer p99 latency - in-region · single-row reads 03 · every run, in order ## Test runs feed. Newest first. Every published test report. Each entry shows the headline metrics inline; the raw JSON lives one click away. loading reports… 04 · how to verify ## All of these are real cargo test runs. Every number on this page is from an actual `cargo test --workspace` run against `live-test-1`, our continuously-deployed managed instance. The build + test script is checked in at `scripts/build-engine.sh`; PDFs of every published run live in our `docs/` folder. If a number on this page is ever wrong, email [support@originchain.ai](mailto:support@originchain.ai) and we will publish a correction with the same prominence as the original. --- # Security controls and compliance | OriginChainDB Canonical source: https://originchaindb.com/trust Sitemap last modified: 2026-09-23T08:33:23.000Z trust center # Security controls and compliance Review implemented controls, certification progress and contractual documents available for your deployment. [Request the security pack](https://originchaindb.com/contact?topic=enterprise) [Security architecture](https://originchaindb.com/security) 01 · status ## Implemented. In progress. Roadmap. implemented - ✓ Encryption - ✓ Tenant isolation - ✓ Audit logging - ✓ Backups (dedicated configurations) - ✓ Point-in-time recovery (dedicated configurations) - ✓ Data residency in progress - … SOC 2 Type 1 (auditor engaged; letter and timeline on request) - … ISO 27001 (certification body engaged; letter and timeline on request) contractual · enterprise - ○ HIPAA Business Associate Agreement - ○ GDPR Data Processing Agreement - ○ Customer-managed encryption keys - ○ Custom retention and quotas 02 · controls, and where they apply ## Implemented today. Scope stated on each. ### Encryption TLS 1.2+ in transit · AES-256 at rest · customer-managed keys on Enterprise. ### Tenant isolation Every dedicated configuration is single-tenant, region-isolated: its own compute, storage, hostname and bearer token. Free and Starter run on shared, scale-to-zero infrastructure. ### Audit logging per-request audit log on every configuration — who called what, when, with which key. ### Backups Dedicated configurations: encrypted nightly snapshots, included; kept inside the instance's region. ### Point-in-time recovery Dedicated configurations: opt-in; restores to a 5 s boundary with retention up to 35 days. Free and Starter do not include point-in-time recovery. ### Data residency Data, transaction logs, backups and the natural-language compile step stay in the region you choose. 03 · documents ## Everything a reviewer asks for first. [Security architecture → Tenancy model, credential handling, what the audit log records, retention, single-row CAS.](https://originchaindb.com/security)[Privacy policy → What we collect about you as a customer and how it is used.](https://originchaindb.com/privacy)[Data processing agreement → Roles, sub-processing and transfer terms under GDPR.](https://originchaindb.com/privacy#dpa)[Terms of service → The governing document, including the availability commitments.](https://originchaindb.com/terms)[Service levels → Availability, recovery and support commitments per configuration.](https://originchaindb.com/sla)[Service status → Live probe results per region.](https://originchaindb.com/status) 04 · responsible disclosure ## Found something? Tell us first. Report vulnerabilities to [security@originchain.ai](mailto:security@originchain.ai) with steps to reproduce. We acknowledge every report, keep you informed until it is resolved, and will not pursue action against research conducted in good faith that respects customer data. in scope - → originchaindb.com (originchain.ai redirects to it), app.originchain.ai and api.originchain.ai - → The published SDKs (TypeScript, Python, Go) - → Your own instance — never another customer's not in scope Denial of service, social engineering, physical attacks, and findings that require access to another tenant's data. ## Security questionnaire in your inbox? Forward it. We answer standard questionnaires and walk your security team through the architecture on a call. [Request the security pack](https://originchaindb.com/contact?topic=enterprise) [Book a technical walkthrough](https://originchaindb.com/contact?topic=walkthrough) --- # OriginChainDB and Aerospike Canonical source: https://originchaindb.com/vs/aerospike Sitemap last modified: 2026-09-23T08:33:23.000Z vs · aerospike # OriginChainDB and Aerospike Compare OriginChainDB and Aerospike across data models, query capabilities, deployment and workload fit. honest framing ## Two AI-database stories. Different shapes. Aerospike's pitch is real-time at hyperscale - 15 years of tail-latency engineering, Hybrid Memory Architecture, ACID at RF2, 256-node clusters, customers like Criteo, Wayfair, Flipkart, DraftKings. It's the right call for ad-tech auctions and instant payments. OriginChainDB's pitch is AI-native primitives - `/ask` (NL → multi-shape query plan), `/watch` (reactive change-stream subscriptions), BYOK LLM (your provider, your audit, your bill), and atomic cross-shape writes that don't need a coordination layer. migration path Both products expose a key/value access model, which makes per-namespace migrations from Aerospike to OriginChainDB feasible - not a wholesale rewrite. If you're considering OriginChainDB for new AI-native workloads while keeping Aerospike for hot ad-tech, that split is sensible. capability matrix | Capability | Aerospike | OriginChainDB | | --- | --- | --- | | Tagline | The real-time database for AI | The AI-native database | | Founded | 2009 | 2026 | | Key-value | Yes - core for 15 years | Yes - substrate | | Document | Collection data types | Via SQL schemas + JSON values | | Graph | Gremlin / Apache TinkerPop | REST + JSON; 9 algorithms shipped | | SQL | Supported across models | Full operational SQL surface | | Vector | Self-healing HNSW | HNSW + sparse + PQ scaffold; 4 metrics | | Natural-language /ask | - | Cost-walker informed, deterministic dispatch | | Reactive /watch | - (CDC via XDR, not query-side) | Live change stream on every shape | | BYOK LLM | - | OpenAI · Anthropic · Gemini · Groq | | Atomic cross-shape writes | ACID at RF2; multi-model scope unclear | One atomic write spans every shape | | Cross-datacenter replication | XDR - multi-DC | Single-region until 1.x | | Connector ecosystem | Spark, Kafka, Pulsar, Elasticsearch, Trino | CSV / NDJSON, webhooks, Postgres / MySQL sync | | Server-side functions / UDFs | Yes | - | | Deployment models | Self-managed · Managed Service · Cloud | Managed SaaS only | | Pricing model | By unique production data volume | Per-configuration RPS + storage caps; structured 429/402 | | Editions | Community (8 nodes, 2.5 TB) · Enterprise · Cloud | Build your configuration · Enterprise | | 2026 recognition | NoSQL Solution of the Year - Data Breakthrough Awards | (new entrant) | pick aerospike when - →Real-time bidding, ad-serving, ad-tech at hyperscale - 15 years of tail-latency engineering. - →Instant payments, intra-day trading - proven enterprise SLAs. - →Multi-datacenter deployments - XDR is real and battle-tested. - →Heavy connector / data-pipeline integration (Kafka, Spark, Trino) is the day-one requirement. - →You need server-side UDFs near data. pick originchaindb when - →Workload is AI-native - RAG, agents, NL queries, reactive personalisation. - →You want one substrate that gives atomic cross-shape writes by construction, not by best-effort. - →/ask, /watch, and BYOK LLM matter - the LLM contract is yours, not the platform's. - →You prefer per-shape SLA transparency (structured 429 / 402) over a single bill-by-data-volume number. - →You want managed SaaS without a sales-cycle, with a self-serve signup. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and Milvus Canonical source: https://originchaindb.com/vs/milvus Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and Milvus Compare OriginChainDB and Milvus across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. The most useful frame for this comparison is scope. Milvus is a vector engine - a serious one, designed for the regime where vectors dominate the workload and the team is willing to run a multi-component cluster to keep up with billion-scale similarity search. OriginChainDB is a database that happens to do vectors as one of five shapes. If you are at billion scale and vectors are the workload, Milvus is built for that. If you are in the 1M - 100M range and the embedding always travels with rows, full-text, and graph relations, the multi-shape managed engine is usually the cleaner answer. choose Milvus if - + You are at billion-vector scale, or planning to be - peak vector throughput is the headline number you optimise for. - + GPU acceleration and index variety (HNSW, IVF, DiskANN, ScaNN, GPU_IVF_FLAT) actively shape your retrieval design. - + Your team has the muscle memory to run a clustered system with several roles - coordinator, query node, data node, index node. - + Vectors dominate the workload, and the surrounding rows live happily in a separate primary database. choose OriginChainDB if - + You operate in the 1M - 100M vector range, where atomic multi-shape writes matter more than peak vector throughput. - + Embeddings travel with structured rows, full-text content, and graph edges that need real queries - not just metadata filters. - + You would rather not run a vector cluster alongside a row store, a search engine, and the sync code that ties them together. - + You want a managed database - single-tenant on a dedicated configuration - and a natural-language endpoint without bolting an LLM service alongside. 02 · where milvus wins ## A vector engine designed for hyperscale. Milvus has spent years targeting the hardest end of the vector workload - billion-vector indexes, multi-tenant clusters, retrieval pipelines that feed search and recommendation surfaces at internet scale. The codebase is open source, the project is active, and the design has been pushed by users who genuinely have those problems. If you are operating at that scale, the engineering investment behind Milvus is real and visible in the product. The index catalogue is one of the broadest in the space. HNSW for low-latency recall, IVF families for memory-bounded indexes, DiskANN for indexes that exceed memory, ScaNN for quantised retrieval, and GPU-resident variants for teams that want to spend GPUs on similarity search. Picking between these is a real operational decision, but the option of picking is itself valuable when your workload genuinely sits at the edge of one of those regimes. The clustered architecture - coordinator, query node, data node, index node, message queue - is heavyweight, but it is heavyweight on purpose. Decoupling those concerns is what lets Milvus scale write ingestion, index building, and query serving independently, which is exactly the right shape for the workloads it was designed for. For teams that have the people to run it (or that pay Zilliz Cloud to run it for them), Milvus is the right tool for the right job. 03 · where originchaindb is different ## A database, not a vector layer. One managed engine, five shapes. The scope difference is the whole story. Milvus is the vector layer of an AI stack - you still run a primary database for rows, often a search engine for full-text, sometimes a graph store, and the application code that keeps them all consistent. OriginChainDB is the database. Rows, secondary indexes, vector embeddings, HNSW graphs, BM25 full-text postings, and graph edges all live in one managed engine. The query engine compiles SQL, vector top-k, BM25 search, graph traversal, and natural-language questions to the same plan tree, and the same executor runs them. The HNSW index has two operating points worth naming concretely: the default `high_recall` mode hits recall@10 = 0.96 at 100k vectors with p99 around 109 ms, and a `fast` mode trades recall for latency, running p99 around 37 ms at recall@10 ≈ 0.69 on the same dataset. For scale, the IVF-PQ index compresses each vector so a large corpus still fits one box; we have not published a measured result at 100M scale. CPU-only, f32 SIMD kernels for cosine, dot, and L2. We are not chasing peak vector throughput at the high end - we are chasing atomicity, predictability, and managed simplicity in the regime where most production AI applications actually live. The structured side is a real query engine, not attribute filtering on a vector record. Full SQL - JOIN, GROUP BY, HAVING, OUTER, LIMIT - runs against the same engine that holds the embedding. Hybrid retrieval - "documents matching this filter, ranked by vector similarity, scored against this BM25 phrase, joined to the author graph, restricted to the last week" - is one statement, one round-trip, one consistent snapshot. With a vector layer plus a primary plus a search engine, that same query is a multi-engine join you write yourself, and the consistency story is whatever your sync code happens to guarantee. Natural language is part of the same surface. `/v1/ask` compiles an English question to the same plan AST as a hand-written query - same cost model, same EXPLAIN output, same per-node statistics. The model emits a plan; the executor runs it. There is no LLM in the hot path and no second service alongside the database. 04 · the dual-write problem ## Why "row + embedding in one atomic commit" matters. The standard architecture for an AI application with Milvus is dual-write, often triple-write: insert the row in your primary database, embed the content, write the vector to Milvus, optionally push a full-text posting to Elastic or OpenSearch, hope nothing crashed in between. Most teams paper over the gap with idempotency keys, retry queues, and reconciliation jobs that scan for orphans across three stores. It works most of the time, and the failure modes are usually invisible until a user reports a search hit that returns no document. OriginChainDB folds the entire derived state into the write path. The row, the embedding, every secondary index, the BM25 postings, every forward and reverse edge - all of them are part of the same atomic write, committing atomically and durably. Partial writes are dropped on recovery, so there is no half-written state to clean up. That property is verified at runtime: a fault-injection harness deliberately crashes the database at multiple boundaries, asserting recovered state equals a prefix of the op stream every time. For applications where the embedding is the only piece of state that matters - a recommender that does not need to know about the document, a similarity index over an immutable corpus at billion scale - the dual-write story is fine and a vector engine like Milvus is a clean fit. For applications where deleting a document has to also delete its embedding, where a row update has to invalidate a stale vector, or where retrieval has to combine vector similarity with a JOIN, BM25 score, or graph hop, one managed engine is the cleaner answer - particularly in the 1M - 100M vector range where atomic multi-shape ops matter more than peak similarity throughput. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | Milvus | OriginChainDB | | --- | --- | --- | | Primary use case | Hyperscale vector similarity | Multi-shape DB - rows + vectors + FTS + graph + NL | | Target scale | Billions of vectors | 10M+ per instance via IVF-PQ; 1M-100M typical | | Index variety | HNSW, IVF, DiskANN, ScaNN, GPU variants | HNSW + IVF + IVF-PQ (Jegou 2011) + binary quant + PQ + sparse + brute | | GPU acceleration | Yes (GPU index families) | CPU SIMD only, no GPU dependency | | Architecture | Clustered - coordinator + query/data/index nodes | Single-binary database; single-tenant on dedicated configurations | | Structured filters | Scalar fields with attribute filtering | Real columns, indexes, JOINs (up to 32 tables), GROUP BY, HAVING, FK + CHECK, materialized views | | Full-text search | Recent BM25 sparse-vector support | Native BM25 + phrase + 18-lang stemming + 9-lang lemmas + ICU, atomic with rows | | Graph traversal | External - your relational DB | Native fwd / rev edges + PageRank + LPA + Node2Vec + GraphSAGE + Cypher v3 | | Atomicity row + embedding | Application-level dual-write | One atomic commit, durable | | Natural-language query | External - your LLM layer | /v1/ask endpoint, plan-bound | | Hosting model | Self-host or Zilliz Cloud | Managed-only; shared on Free and Starter, dedicated infrastructure on dedicated configurations | | Operations footprint | Multi-component cluster | One service that replaces row-store + vector + FTS + graph | 06 · operations ## Two different operational stories. Milvus is a clustered system. Coordinator, query nodes, data nodes, index nodes, plus a message queue and an object store underneath. Each role can scale independently, which is the point at billion scale - write ingestion, index building, and query serving have very different resource profiles, and decoupling them is what lets the system reach the regime it was designed for. The trade-off is that running it well is a real engineering commitment. Most teams that do not need that scale outsource the operational burden to Zilliz Cloud. OriginChainDB is managed-only by design. A dedicated configuration gives each customer a single-tenant database in a region of their choice, with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. We have intentionally not designed for billion-scale vector throughput; in exchange, we have made the operational footprint a single managed instance with no clustered roles to balance. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and MongoDB Canonical source: https://originchaindb.com/vs/mongodb Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and MongoDB Compare OriginChainDB and MongoDB across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. MongoDB has earned its place. The document model is a real fit for nested, evolving, application-defined shapes - content management, IoT telemetry, configuration, anywhere relational normalisation feels like a tax. Atlas wraps it in a mature managed cloud, and Atlas Vector Search has put a credible vector capability inside the same product. The interesting question is not which database is "better" - it is whether your application is shaped like a document store that occasionally needs vectors, or like an AI workload where vectors, full-text, graph, and rows have to commit consistently together. The answer to that question decides the database. choose MongoDB if - + Workload is document-shaped - nested data, evolving fields, per-record schema flexibility is genuinely useful. - + You are already a MongoDB shop with drivers, ORMs, and operational muscle memory in place. - + Atlas's wider surface - Search, Charts, App Services, Triggers - earns its keep alongside the database. - + Vector similarity is one of several access patterns and you are comfortable with Atlas Vector Search's separate-index model. choose OriginChainDB if - + AI features are the workload - embeddings, hybrid search, graph context, natural language are equal citizens to rows. - + Typed schema is acceptable, even desirable - TOML manifests give the planner the type information AI features depend on. - + Rows, embeddings, full-text postings, and graph edges have to commit atomically in one round-trip. - + Single-tenant infrastructure matters: on a dedicated configuration you do not share cluster resources with other customers. 02 · where mongodb wins ## Fifteen years of document maturity, plus a managed cloud that actually works. MongoDB's document model is a genuine fit for nested, application-shaped data. A user profile with embedded preferences, an order with line items and shipping events, a CMS document with arbitrarily deep blocks - these things model badly as third-normal-form rows and beautifully as BSON documents. The drivers are mature in every language anyone uses, the aggregation pipeline is more capable than most relational engines give it credit for, and the schemaless flexibility is genuinely useful when fields evolve faster than migrations. Atlas is the second pillar. It is a mature managed cloud - multi-region replica sets, sharded clusters, point-in-time backup, online resharding, encryption everywhere, granular RBAC, a credible compliance posture. The adjacent products are real: Atlas Search puts Lucene-backed full-text on top of the same documents, Atlas Vector Search added HNSW indexes in 2023, App Services and Triggers cover server-side logic, and Charts handles in-product analytics. For an existing MongoDB shop, the marginal cost of adding vector search to documents you are already storing is unusually low. The ecosystem effect compounds. Mongoose, Prisma, Beanie, every popular ODM has a polished MongoDB path. Every CDC pipeline knows how to read the oplog. Every BI tool has a connector. For workloads where the document model fits and the wider Atlas surface earns its keep, picking MongoDB is often the right call and a low-risk one. 03 · where originchaindb is different ## Typed schema. Atomic multi-shape. One managed engine. OriginChainDB takes the opposite bet on schema. Tables, indexes, vector fields, full-text fields, and graph edges are declared in TOML manifests with explicit types. That is more friction up front than schemaless documents, but it is a deliberate trade-off: AI features benefit enormously from typed columns, because the planner can reason about cardinality, push predicates below vector distance computation, and choose the right index without reading every document to find out what fields it has. The schema is also a contract - a row update can invalidate a stale embedding because the system knows which column generated it. Underneath that, OriginChainDB is a single managed engine. Rows, secondary indexes, vector embeddings, HNSW graphs, BM25 full-text postings, and graph edges all live in that engine. The query engine compiles SQL, vector top-k, BM25 search, graph traversal, and natural-language questions to the same plan tree. HNSW has two operating points worth naming: the default `high_recall` mode hits recall@10 = 0.96 at 100k vectors with p99 around 109 ms, and a `fast` mode runs p99 around 37 ms at recall@10 ≈ 0.69. For scale, the IVF-PQ index compresses each vector so a large corpus still fits one box; we have not published a measured result at 100M scale. On a dedicated configuration, each customer gets a single-tenant database in a region of their choice - its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Atlas pools many customers onto shared cluster infrastructure (with logical isolation), which is the right trade-off for the price points and workloads they target; OriginChainDB's dedicated configurations make the opposite trade-off, while its Free and Starter plans share infrastructure too. If compute and storage that no other customer shares matter for your compliance or performance budget, that is a meaningful difference. Graph is native, not a pipeline. A graph traversal in OriginChainDB uses real forward and reverse edge indexes and shortest-path algorithms (BFS, Dijkstra) - not `$graphLookup` over documents. Natural language is part of the same surface: `/v1/ask` compiles an English question to the same plan AST as a hand-written query. The model emits a plan; the executor runs it. There is no LLM on the hot path and no second service to deploy. 04 · the atomicity gap ## What "doc + embedding in one atomic commit" actually buys you. Atlas Vector Search is implemented as a separate index over a document collection - a Lucene-backed HNSW index that is updated asynchronously from the primary collection's writes. That is the right architecture for the document model: it lets you bolt vectors onto existing collections without changing the storage layout. The trade-off is that a write to the document and the corresponding update to the vector index are not part of the same atomic commit. Most of the time the lag is small and invisible, but during failover, heavy load, or index rebuilds, there is a window where a document has been updated and its embedding has not. OriginChainDB folds the embedding into the same write as the row. A single insert writes the row, every secondary index, every graph edge update, the full-text postings, and the vector - all as one atomic write, committing atomically and durably. Partial writes are dropped on recovery; there is no half-written state where the row exists but the vector does not. Recovery correctness is verified at runtime by a fault-injection harness that crashes the database at multiple boundaries, asserting recovered state equals a prefix of the op stream every time. For applications where the embedding is a derived view that can lag the document briefly - recommender systems over an immutable corpus, similarity search where freshness in seconds is fine - Atlas Vector Search's separate-index model is a clean fit. For applications where deleting a document has to also delete its embedding atomically, where a row update has to invalidate a stale vector synchronously, or where retrieval has to combine vector similarity with a row-level filter and a graph hop in one snapshot, one managed engine is the cleaner answer. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | MongoDB | OriginChainDB | | --- | --- | --- | | Data model | Schemaless documents (BSON) | Typed schema via TOML manifests | | Tenancy model | Shared cluster (Atlas) by default | Shared on Free and Starter; single-tenant on dedicated configurations | | Vector search | Atlas Vector Search (Lucene HNSW) | Native HNSW + f32 SIMD | | Full-text | Atlas Search (Lucene-backed) | Native BM25 + phrase + stemming | | Graph traversal | $graphLookup / $lookup pipelines | Native fwd / rev edges + Dijkstra | | Atomicity row + embedding | Vector index updated separately from doc | Doc + index + embedding + posting + edge in ONE atomic commit | | Natural-language query | Bring-your-own LLM layer | /v1/ask endpoint, plan-bound | | Transactions | Multi-document, replica-set scoped | Multi-shape, single atomic commit, durable before ack | | Schema evolution | Implicit, application-managed | Explicit migrations against typed manifests | | Replication | Replica sets + sharded clusters | Async standby + fenced automatic failover | | Pricing shape | Cluster tier + storage + transfer | Free, Starter, or a dedicated compute configuration + flat add-ons | | Operations footprint | Atlas + adjacent services (Search, etc.) | One service that replaces row-store + vector + FTS + graph | 06 · operations ## Two different operational stories. Atlas's operational story is well-trodden. Pick a cluster tier, choose a region (or several), enable the features you need - Search, Vector Search, App Services, Triggers, Charts - and the platform handles backups, replication, sharding, and upgrades. For a team with existing MongoDB skills, the operational ceiling is high and the failure modes are well understood. The trade-off is that running multiple Atlas products alongside the database (Search nodes, Vector Search workloads, App Services functions) compounds into a footprint with several dashboards, several pricing dimensions, and several places where state can lag. OriginChainDB replaces several of those pieces with one managed database per region. A dedicated configuration gives each customer a single-tenant instance with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. The trade-off is real: you give up the breadth of Atlas's adjacent product surface in exchange for a database where vector / full-text / graph / NL are first-class, atomic across shapes, and single-tenant on dedicated configurations. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and Neon Canonical source: https://originchaindb.com/vs/neon Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and Neon Compare OriginChainDB and Neon across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. Neon is one of the more interesting things to happen to Postgres in years. Separating compute from storage is not just an architecture choice - it makes branching, point-in-time recovery, and scale-to-zero cheap in a way they have never been on a traditional Postgres deployment. The interesting question for your team is whether your workload is Postgres-shaped, where serverless and branching are the killer features, or whether your workload is AI-shaped, where atomic multi-shape writes and native vector / FTS / graph are the killer features. The answer to that question decides the database. choose Neon if - + Workload is Postgres-shaped - relational primary, occasional vector via pgvector, predictable SQL access patterns. - + Database branching for preview environments is a killer feature for your team's review-app workflow. - + Bursty or low-utilisation traffic where scale-to-zero materially changes the bill. - + You are comfortable layering pgvector and other extensions, and tracking their compatibility with the platform. choose OriginChainDB if - + AI features are the workload - embeddings, hybrid search, graph context, natural language are equal citizens to rows. - + You want vector / full-text / graph as native shapes, not as extensions you have to wire and version-pin. - + Rows, embeddings, full-text postings, and graph edges have to commit atomically in one round-trip. - + Single-tenant compute matters: on a dedicated configuration you do not share storage and compute pools with other customers. 02 · where neon wins ## Postgres, reshaped for the cloud. Neon's architectural bet is that Postgres's storage layer should live somewhere that scales independently of the query engine. The result is a Postgres that branches like git - you can fork an entire database in seconds, run a destructive migration on the fork, point a preview environment at it, and throw the branch away when the pull request closes. Teams that adopt that workflow tend to keep it forever; the alternative of seeding a fresh database for every preview environment, or sharing a stale staging database across the entire team, is genuinely worse. Scale-to-zero is the second big win. Most application databases sit idle most of the time, and most managed Postgres vendors charge you for that idle. Neon spins compute down when there are no queries and brings it back fast on the next request, which materially changes the bill for low-utilisation projects, hobby workloads, and test environments. Combined with the consumption-based storage pricing, the cost shape is much closer to "what you actually used" than the traditional always-on instance. Underneath all of it, you still get Postgres. Every ORM works. Every migration tool works. pgvector is available as an extension, tsvector / GIN are there for full-text, recursive CTE handles a fair amount of graph work. If your application is dominantly relational and the AI features are additive, Neon gives you an unusually nice cloud-native Postgres without making you reach outside the ecosystem. 03 · where originchaindb is different ## Different engine. Atomic multi-shape from day one. OriginChainDB is not Postgres-on-object-storage. It is a different engine - a single managed engine - designed from the start to hold rows, secondary indexes, vector embeddings, HNSW graphs, BM25 full-text postings, and graph edges in the same place. The query engine compiles SQL, vector top-k, BM25 search, graph traversal, and natural-language questions to the same plan tree. There is no extension stack to wire and no version drift between pgvector, ParadeDB, or Apache AGE - the AI shapes are not extensions, they are first-class. HNSW has two operating points worth naming: `high_recall` at recall@10 = 0.96 with p99 around 109 ms on 100k vectors, and `fast` at recall@10 ≈ 0.69 with p99 around 37 ms. For scale, the IVF-PQ index compresses each vector so a large corpus still fits one box; we have not published a measured result at 100M scale. The consequence is atomicity that crosses shapes. A single insert writes the row, every secondary index entry, every forward and reverse edge, the BM25 postings, and the vector embedding in one atomic write. That write commits atomically and replicates to the standby as one unit. With Neon's separation of compute and storage, Postgres's transactional guarantees still hold within Postgres, but the ops shape carries new tunables - page-server caching, autosuspend behaviour, branch retention, compute-size scaling. OriginChainDB has fewer moving parts because there is no separation-of-storage layer to configure; the ops surface is simpler by construction. On a dedicated configuration, each customer gets a single-tenant database in a region of their choice, with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Neon shares compute pools and a shared object-store backing, which is the right trade-off for branching and scale-to-zero; OriginChainDB's dedicated configurations make the opposite trade-off, optimising for predictable per-tenant performance and isolation, while its Free and Starter plans, like Neon, run on shared, scale-to-zero infrastructure. Natural language is part of the same surface. `/v1/ask` compiles an English question to the same plan AST as a hand-written query - same cost model, same EXPLAIN output, same per-node statistics. The model emits a plan; the executor runs it. There is no LLM on the hot path, no token-priced query layer to budget for, and no second service to deploy alongside the database. 04 · the atomicity gap ## What "one INSERT, one atomic commit" actually buys you. The standard architecture for an AI feature on Neon is a Postgres transaction that touches a row, a tsvector column, and a pgvector column. Inside one transaction, that is genuinely atomic - Postgres's MVCC keeps the writes consistent. The seam shows up when the AI surface grows. Add a separate full-text engine for typo-tolerant retrieval, a graph store for context that does not fit recursive CTE, or an out-of-process embedding worker, and the consistency story splits across services. Most teams paper over the gap with idempotency keys and reconciliation jobs, all of which work until they don't. OriginChainDB folds the entire derived state into the write path. The row, the embedding, the full-text postings, every edge update - all of them are part of the same atomic write, which commits as one unit. Partial writes are dropped on recovery, so there is no half-written state to clean up. That property is verified at runtime: a fault-injection harness deliberately crashes the database at multiple boundaries, and recovery is asserted to equal a prefix of the op stream every time. We run it for a million deterministic iterations on every build. Reads compose the same way. A query can filter on structured columns, rank by vector similarity, intersect with a BM25 search, and join across a graph edge - in one round trip, against one consistent snapshot. With pgvector + tsvector inside one Postgres statement, you can get a long way; once a fourth shape (or a non-Postgres engine) enters the picture, you are stitching results in application code. OriginChainDB exists because that boundary is exactly where AI applications keep getting bitten. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | Neon | OriginChainDB | | --- | --- | --- | | Core engine | Postgres with separated compute/storage | Single managed engine, multi-shape | | Tenancy model | Multi-tenant compute on shared storage | Shared on Free and Starter; single-tenant on dedicated configurations | | Branching | Copy-on-write database branches | Snapshot-based clones; no per-branch isolation | | Scale-to-zero | Yes, fast cold start | Free and Starter scale to zero; dedicated configurations are always on | | Vector search | pgvector extension | Native HNSW + IVF + IVF-PQ + binary/PQ quantization + sparse | | Full-text | tsvector / GIN built-in | Native BM25 + phrase + 18-lang stemming + 9-lang lemmas + ICU | | Graph traversal | Recursive CTE | Native fwd / rev edges + PageRank + LPA + Node2Vec + GraphSAGE + Cypher v3 | | Atomicity across shapes | Per-row within a Postgres transaction | Row + index + embedding + posting + edge in ONE atomic commit | | Natural-language query | Bring-your-own LLM layer | /v1/ask endpoint, plan-bound | | Materialized views | Manual + Postgres extensions | On-demand install / refresh | | FK + CHECK | Standard Postgres | FK (NoAction / Restrict / SetNull) + CHECK with 3-valued logic | | Postgres data import | Native (it is Postgres) | POST /ingest/postgres/sync pulls from existing Postgres source | | Replication | Postgres replicas + S3-backed storage | Async standby + fenced automatic failover; multi-writer in development | | Pricing shape | Compute hours + storage usage | Free, Starter, or a dedicated compute configuration + flat add-ons | | Operations footprint | Managed Postgres with serverless ops | One service that replaces row-store + vector + FTS + graph | 06 · operations ## Two different operational stories. Neon's operational story is unusual for a Postgres vendor: the database can sleep when nobody is using it, branch in seconds, and bill in something close to "what you actually consumed." That is genuinely useful, especially for preview environments and projects where the median traffic shape is bursty. The trade-off is that there are knobs unique to the architecture - autosuspend windows, page-server caching, branch retention, compute-size sizing - and Postgres extension compatibility lives within whatever the platform supports for that release. OriginChainDB is managed-only by design. A dedicated configuration gives each customer a single-tenant database in a region of their choice, with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. The trade-off is real: on a dedicated configuration you give up scale-to-zero and copy-on-write branching in exchange for a database where vector / full-text / graph / NL are first-class, atomic across shapes, and single-tenant. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and Pinecone Canonical source: https://originchaindb.com/vs/pinecone Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and Pinecone Compare OriginChainDB and Pinecone across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. The interesting question is not "which is faster" - both products can serve sub-100 ms top-k vector queries, and the differences in raw throughput depend more on workload shape than on benchmark folklore. The interesting question is what is on the other side of the embedding. If it is rows, full-text, and relations that need to stay consistent with the embedding, you probably want one database. If it is purely a vector index with a primary store somewhere else, you probably want a vector specialist. choose Pinecone if - + Workload is pure vector - similarity search at scale, with structured filters that fit on the vector record itself. - + You already have a relational primary you trust, and you are happy with the dual-write story. - + Index size or query rate dominates everything else, and you want a vendor whose entire roadmap is similarity search. - + You can absorb a small consistency window between your row store and your vector store. choose OriginChainDB if - + Vectors travel with rows - every embedding has structured columns, full-text content, or graph relations alongside it. - + Hybrid retrieval (vector + filter + BM25 + graph) needs to run as one query against one consistent state. - + You would rather not run two databases for one feature, or write reconciliation jobs to keep them aligned. - + You want a managed database - single-tenant on a dedicated configuration - and a natural-language endpoint without bolting an LLM service alongside. 02 · where pinecone wins ## A vector specialist that has earned its reputation. Pinecone has been one of the load-bearing pieces of the modern AI stack since the early LangChain era. The team has spent years tuning a similarity-search engine that handles billion-vector indexes, namespace isolation, sparse-dense hybrid retrieval, and the operational realities of running vector search at scale. The serverless tier in particular is a strong fit for workloads with bursty query patterns, where the alternative is paying for idle pod capacity. The brand and ecosystem matter too. Most retrieval-augmented-generation tutorials use Pinecone in their first example, every popular framework (LangChain, LlamaIndex, Haystack, Semantic Kernel) has a polished integration, and a lot of senior engineers have a working mental model of how it behaves. If your bottleneck is similarity search and your team is already shipping with it, the marginal cost of staying is low and the path is well lit. Pinecone is also genuinely opinionated about doing one thing well. The metadata story is intentionally minimal - fields on the vector record, predicate filtering, namespace partitioning - because the product is not trying to be your relational database. If your application can express its filters that way, the simplicity is a feature, not a limitation. 03 · where originchaindb is different ## Vector is one shape. The same store holds the rest. OriginChainDB is built around a single managed engine. Vectors live alongside the rows they describe, the full-text postings that share their content, and the graph edges that connect them. An HNSW graph backs vector top-k with f32 SIMD kernels for cosine, dot, and L2 distance, and the index has two operating points worth naming concretely: the default `high_recall` mode hits recall@10 = 0.96 at 100k vectors with p99 around 109 ms, and a `fast` mode trades recall for latency, running p99 around 37 ms at recall@10 ≈ 0.69. For scale, the IVF-PQ index compresses each vector so a large corpus still fits one box; we have not published a measured result at 100M scale. You pick the operating point per workload. Because the embedding lives next to the row, structured filters are real columns rather than metadata appended to the vector record. You write `SELECT` with predicates, joins, group-by, and ordering - all in the same query that produces the top-k. The cost model picks between full scan and index scan from query statistics, and SIMD predicates run before vector distance is computed, so the planner can prune work without juggling two engines. Hybrid retrieval is a single plan tree. A query that wants "top-k by vector similarity, restricted to documents posted in the last week, with a BM25 boost from a keyword phrase, joined to the author graph" is one statement against one managed engine. With a vector specialist you would do the BM25 in your search engine, the row filter in your relational database, the graph hop in a third store, and stitch the result in application code - every join across engines is a network hop and a consistency assumption. Natural language is part of the same surface. `/v1/ask` compiles an English question to the same plan AST as a hand-written query. The model emits a plan; the executor runs it. There is no LLM in the hot path, no token-priced query layer to budget for, and no second service alongside the database. 04 · the dual-write problem ## Why "row + embedding in one atomic commit" matters. The standard architecture for an AI application with Pinecone is dual-write: insert the row in your relational database, embed the content, write the vector to Pinecone, hope nothing crashed in between. Most teams paper over the gap with idempotency keys, retry queues, and a reconciliation job that scans for orphaned rows or orphaned vectors. It works most of the time, and the failure modes are usually invisible until a user reports a search hit that returns no document. OriginChainDB folds the embedding into the same write as the row. A single insert writes the row, every secondary index, every graph edge update, the full-text postings, and the vector - all as one atomic write, committing atomically and durably. Partial writes are dropped on recovery; there is no half-written state where the row exists but the vector does not. Recovery correctness is verified at runtime by a fault-injection harness that crashes the database at multiple boundaries, asserting recovered state equals a prefix of the op stream every time. For applications where the embedding is the only piece of state that matters - a recommender that does not need to know about the document, a similarity index over an immutable corpus - the dual-write story is fine and Pinecone is a clean fit. For applications where deleting a document has to also delete its embedding, where a row update has to invalidate a stale vector, or where retrieval has to combine vector similarity with a row-level filter, one managed engine is the cleaner answer. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | Pinecone | OriginChainDB | | --- | --- | --- | | Primary use case | Pure vector similarity at scale | Multi-shape DB - rows + vectors + FTS + graph + NL | | Structured filters | Metadata fields on the vector record | Real columns, indexes, JOINs, and aggregates | | Full-text search | Sparse-dense hybrid (Pinecone hybrid) | Native BM25 + phrase + 18-lang stemming + 9-lang lemmas + ICU, atomic with rows | | Graph traversal | External - your relational DB | Native fwd / rev edges + PageRank + LPA + Node2Vec + GraphSAGE + Cypher v3 | | Atomicity row + embedding | Application-level dual-write | One atomic commit, durable | | Natural-language query | External - your LLM layer | /v1/ask endpoint, plan-bound | | Tenancy model | Multi-tenant serverless / pod | Shared on Free and Starter; single-tenant on dedicated configurations | | Region isolation | Region-pinned | Region-pinned; dedicated infrastructure on dedicated configurations | | Vector index | Proprietary - well-tuned at scale | HNSW + IVF + IVF-PQ (64× memory) + binary quant (32× memory) + PQ + sparse | | Pricing shape | Pod-based or serverless usage | Free, Starter, or a dedicated compute configuration + flat add-ons | | Operations footprint | One service to operate | One service that replaces row-store + vector + FTS + graph | 06 · operations ## Two different operational stories. Pinecone's operational story is "we run the vector database, you don't." That is genuinely useful. You provision an index, choose a region and a pod size or the serverless tier, and ship. The downside is that it is one piece of a multi-database stack - to ship a typical AI feature you also operate a relational primary, often a search engine, sometimes a graph store, and the application code that keeps them all in sync. Each one has its own dashboard, its own credentials, its own backup schedule, its own failure modes. OriginChainDB replaces several of those pieces with one managed database per region. A dedicated configuration gives each customer a single-tenant instance with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and PostgreSQL Canonical source: https://originchaindb.com/vs/postgres Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and PostgreSQL Compare OriginChainDB and PostgreSQL across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. We are not going to argue you should rip out Postgres. For most relational applications, Postgres is the right answer and will be for a long time. The question worth asking is whether your workload is shaped like a relational application that occasionally needs an embedding, or like an AI application that occasionally needs a row. The answer to that question decides the database. choose Postgres if - + Workload is predominantly relational - orders, users, billing, the long tail of an OLTP application. - + Your team already runs Postgres at scale and has the muscle memory for tuning, replication, and extensions. - + You depend on the ecosystem - every ORM, every migration tool, every BI connector speaks Postgres. - + Vector / full-text / graph are secondary concerns you can carry with pgvector, ParadeDB, or Apache AGE. choose OriginChainDB if - + AI features are the workload - embeddings, hybrid search, graph context, natural language are not extras. - + Rows, embeddings, full-text postings, and graph edges have to commit consistently in one round-trip. - + You would rather not operate four data systems (relational + vector DB + search + graph) and the sync code between them. - + You want a managed database - single-tenant on a dedicated configuration - that ships natural-language query as a first-class endpoint, not as a bolt-on. 02 · where postgres wins ## Twenty-five years of relational maturity. Postgres has earned its place as the default primary database. Its planner is one of the most sophisticated open-source query optimisers ever written. MVCC, write-ahead logging, streaming and logical replication, partitioning, foreign data wrappers, and the procedural-language story (PL/pgSQL, PL/Python, PL/v8) cover an enormous surface of relational workloads. The community has shipped hundreds of extensions - pgvector for embeddings, PostGIS for geospatial, TimescaleDB for time series, ParadeDB for full-text, Apache AGE for graphs, Citus for sharding - meaning you can usually find a starting point for whatever shape your data takes. The ecosystem effect is real and it matters. Every modern ORM (Prisma, SQLAlchemy, ActiveRecord, Drizzle, GORM) has a first-class Postgres path that has been tuned for years. Every managed cloud (RDS, Cloud SQL, Aurora, every regional Postgres-as-a-service vendor) ships a battle-tested Postgres. Every BI tool, every observability vendor, every CDC pipeline knows how to read it. For workloads where the dominant question is "what does my application's relational state look like right now," that ecosystem is hard to beat. And Postgres is genuinely capable in adjacent shapes. pgvector now supports HNSW indexes; tsvector + GIN gives you respectable full-text search; recursive CTE and Apache AGE handle a lot of graph workloads; logical replication + Debezium gives you a reasonable change-feed. None of these are toys, and for many teams they are the right call: keep one database, add one extension, ship. 03 · where originchaindb is different ## One managed engine. Five shapes. One atomic commit per write. OriginChainDB is built around a single managed engine. Rows, secondary indexes, vector embeddings, HNSW graphs, BM25 full-text postings, and graph edges all live in that engine. The query engine compiles SQL, vector top-k, BM25 search, graph traversal, and natural-language questions to the same plan tree, and the same executor runs them. There are no extensions to wire together because there is no second engine to wire to - every shape is a first-class capability against the same data. The consequence is atomicity that crosses shapes. A single insert writes the row, every secondary index entry, every forward and reverse edge, the BM25 postings, and the vector embedding in one atomic write. That write commits and replicates as one unit. There is no window where the row exists but its embedding does not, no partial state where the full-text posting is half-written, no eventual-consistency drift between your primary and your vector store. Reads compose the same way. A query can filter on structured columns, rank by vector similarity, intersect with a BM25 search, and join across a graph edge - in one round trip, against one consistent snapshot. With Postgres + pgvector + ParadeDB + AGE, the same query is a multi-engine join you write yourself, and the consistency story is whatever your sync code happens to guarantee. Natural language is part of the same surface. `/v1/ask` compiles an English question to the same plan AST as a hand-written query - same cost model, same EXPLAIN output, same per-node statistics. The model emits a plan; the executor runs it. There is no LLM on the hot path, no token-priced query layer to budget for, and no second service to deploy alongside the database. 04 · the atomicity gap ## What "one insert, one atomic commit" actually buys you. In a typical AI stack - Postgres for rows, a vector database for embeddings, an LLM service for answers - the application code holds the consistency story. Insert a document, embed it, write the embedding, hope nothing crashed in between. If it did, you have a row with no embedding, an embedding with no row, or a half-written index. Most teams paper over this with idempotency keys, retry queues, and reconciliation jobs, all of which work until they don't. OriginChainDB folds the entire derived state into the write path. The row, the embedding, the full-text postings, every edge update - all of them are part of the same atomic write, which commits as one unit. Partial writes are dropped on recovery, so there is no half-written state to clean up. That property is verified at runtime: a fault-injection harness deliberately crashes the database at multiple boundaries, and recovery is asserted to equal a prefix of the op stream every time. We run it for a million deterministic iterations on every build. Postgres's transactional guarantees are excellent within the database, but they end at the database boundary. If your embedding lives in Pinecone or your full-text index lives in Elastic, you are back to writing your own two-phase commit, or accepting eventual consistency. OriginChainDB exists because that boundary is exactly the place AI applications keep getting bitten. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | Postgres | OriginChainDB | | --- | --- | --- | | Data model | Heap-and-extension relational | Single managed engine, multi-shape | | Tenancy model | Shared by default, single-tenant by config | Shared on Free and Starter; single-tenant on dedicated configurations | | SQL coverage | Full SQL - JOIN, CTE, window, MVCC | JOIN (up to 32 tables), GROUP BY, OUTER, HAVING, ORDER BY, window functions (incl. LAG/LEAD), correlated + uncorrelated subqueries (IN / EXISTS / scalar), materialized views (on-demand), FK + CHECK | | Vector search | pgvector extension | Native HNSW + IVF + IVF-PQ + binary/PQ quantization + sparse (CSR) | | Full-text | tsvector / GIN, ParadeDB extension | Native BM25 + phrase + 18-lang stemming + 9-lang lemmatization + ICU + geo | | Graph traversal | Recursive CTE, Apache AGE extension | Native fwd / rev edges + Dijkstra + PageRank + betweenness + eigenvector + LPA + Node2Vec + GraphSAGE + Cypher v3 (CALL / FOREACH / MERGE / DELETE) | | Natural-language query | Bring-your-own LLM layer | /v1/ask endpoint, plan-bound | | Atomicity across shapes | Per-table within a transaction | Row + index + embedding + posting + edge in ONE atomic commit | | Replication | Streaming, logical, sync / async | Async standby + fenced automatic failover; multi-writer in development | | Recovery | PITR via WAL archive | Point-in-time recovery, crash-tested | | Operations | Self-hosted or any managed Postgres | Managed-only - no DBA, no extensions to wire | | Postgres data import | — | POST /ingest/postgres/sync pulls from an existing Postgres source | | Ecosystem reach | 25+ years, every ORM, every BI tool | REST + a thin SDK; smaller surface, simpler glue | 06 · operations ## Two different operational stories. Postgres is mature enough that you have real choice on how to run it. Self-host on your own metal, ship to RDS or Cloud SQL, pay for Aurora, run a regional managed Postgres - each option has a thriving ecosystem and well-understood failure modes. If your team already has the muscle memory for tuning `shared_buffers`, planning autovacuum, and reading `pg_stat_statements`, that knowledge transfers cleanly between vendors. OriginChainDB is managed-only by design. A dedicated configuration gives each customer a single-tenant database in a region of their choice, with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. The trade-off is real: you get fewer knobs to turn and a smaller ecosystem, but you also do not need a DBA on call to add vector search. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and Qdrant Canonical source: https://originchaindb.com/vs/qdrant Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and Qdrant Compare OriginChainDB and Qdrant across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. Qdrant and OriginChainDB solve overlapping but different problems. Qdrant is a vector-and-payload engine - vectors with rich JSON metadata, indexed filters, top-k that respects them. OriginChainDB is a multi-shape database - rows, vectors, full-text, graph, and natural-language all compiling to the same plan tree. The interesting question is whether the structured side of your workload is shaped like "metadata on a vector" or shaped like "a database your application reads and writes that also embeds." Both can be right; they are usually not the same workload. choose Qdrant if - + Workload is vector-first - similarity search at scale, with payload filters that fit cleanly on the vector record. - + You want an open-source vector database in Rust that you can self-host or run on a managed cloud. - + Payload indexing - categorical, numeric, geo, full-text on the payload - covers the filter shapes you need. - + You already have a relational primary you trust, and the dual-write story between it and the vector store is acceptable. choose OriginChainDB if - + The embedding always travels with structured row data, full-text content, or graph relations that need real queries. - + Hybrid retrieval (vector + JOIN + BM25 + graph) needs to run as one query against one consistent state. - + You would rather not run two databases for one feature, or write reconciliation jobs to keep them aligned. - + You want a managed database - single-tenant on a dedicated configuration - and a natural-language endpoint without bolting an LLM service alongside. 02 · where qdrant wins ## A Rust-native vector specialist with rich payload filtering. Qdrant is one of the most capable open-source vector databases shipping today. The codebase is in Rust, the design choices read as deliberate, and the operational footprint is small enough that a small team can run it themselves without an operator. The community has grown steadily over the last few years, and Qdrant has become a default option whenever a team needs a self-hostable vector store with predictable behaviour. The payload story is genuinely good. Each vector record carries a JSON payload, and Qdrant supports payload indexes - keyword, integer range, float range, geo, and a full-text payload index - so filters do not degenerate into a post-filter scan. For workloads where the filter shape is "category equals X, price between A and B, location within R kilometres of P" sitting alongside the embedding, Qdrant's filtering is one of the more polished offerings in the space, and it is a deliberate, documented part of the engine. Qdrant is also opinionated about doing one thing well. The product is not trying to be your relational database - it is a vector engine with payloads, full stop. If your application can express its filters that way, the simplicity is a feature: fewer concepts to learn, fewer footguns, and a smaller blast radius for operational mistakes. 03 · where originchaindb is different ## Vector is one shape. The same store holds the rest. OriginChainDB is built around a single managed engine. Vectors live alongside the rows they describe, the full-text postings that share their content, and the graph edges that connect them. An HNSW graph backs vector top-k with f32 SIMD kernels for cosine, dot, and L2 distance. The index has two operating points worth naming concretely: the default `high_recall` mode hits recall@10 = 0.96 at 100k vectors with p99 around 109 ms, and a `fast` mode trades recall for latency, running p99 around 37 ms at recall@10 ≈ 0.69 on the same dataset. For scale, the IVF-PQ index compresses each vector so a large corpus still fits one box; we have not published a measured result at 100M scale. The structured side is a real query engine, not metadata on a record. You write SQL - JOIN across up to 32 tables, GROUP BY, OUTER, HAVING, LIMIT - against the same engine that holds the vector. The cost model picks between full scan and index scan from query statistics; SIMD predicates run before vector distance is computed; aggregates push down. A query that wants "documents matching this filter, ranked by vector similarity, joined to the author table, restricted to the last week" is one statement, not a payload filter on a vector record. BM25 full-text is a first-class shape, not a payload sub-feature. Phrase queries, stemming, and proper BM25 scoring run against the same posting list that the row points at, atomically. Graph is also a first-class shape - forward and reverse edges, Dijkstra, traversal primitives - so retrieval can hop along relationships without leaving the database. With Qdrant you would do the BM25 in a search engine, the JOIN in a relational primary, and the graph hop in a third store; every join across engines is a network hop and a consistency assumption. Natural language is part of the same surface. `/v1/ask` compiles an English question to the same plan AST as a hand-written query - same cost model, same EXPLAIN output, same per-node statistics. The model emits a plan; the executor runs it. There is no LLM in the hot path and no second service alongside the database. 04 · the dual-write problem ## Why "row + embedding in one atomic commit" matters. The standard architecture for an AI application with Qdrant is dual-write: insert the row in your relational database, embed the content, write the vector and payload to Qdrant, hope nothing crashed in between. Most teams paper over the gap with idempotency keys, retry queues, and a reconciliation job that scans for orphaned rows or orphaned vectors. It works most of the time, and the failure modes are usually invisible until a user reports a search hit that returns no document, or a deleted row keeps appearing in similarity results. OriginChainDB folds the embedding into the same write as the row. A single insert writes the row, every secondary index, every graph edge update, the BM25 postings, and the vector - all as one atomic write, committing atomically and durably. Partial writes are dropped on recovery; there is no half-written state where the row exists but the vector does not. Recovery correctness is verified at runtime by a fault-injection harness that crashes the database at multiple boundaries, asserting recovered state equals a prefix of the op stream every time. For applications where the embedding is the only piece of state that matters - a recommender that does not need to know about the document, a similarity index over an immutable corpus - the dual-write story is fine and Qdrant is a clean fit. For applications where deleting a document has to also delete its embedding, where a row update has to invalidate a stale vector, or where retrieval has to combine vector similarity with a row-level JOIN, BM25 score, or graph hop, one managed engine is the cleaner answer. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | Qdrant | OriginChainDB | | --- | --- | --- | | Primary use case | Pure vector similarity with rich payload | Multi-shape DB - rows + vectors + FTS + graph + NL | | Data model | Vectors with JSON payloads | Single managed engine, multi-shape | | Structured filters | Payload filters + payload indexes | Real columns, indexes, JOINs, GROUP BY, HAVING | | Full-text search | Full-text payload index (substring) | Native BM25 + phrase + 18-lang stemming + 9-lang lemmas + ICU, atomic with rows | | Graph traversal | External - your relational DB | Native fwd / rev edges + PageRank + LPA + Node2Vec + GraphSAGE + Cypher v3 | | Vector index | HNSW (Rust, well tuned) | HNSW + IVF + IVF-PQ (64× memory) + binary quant (32× memory) + PQ + sparse | | Atomicity row + embedding | Application-level dual-write | One atomic commit, durable | | Natural-language query | External - your LLM layer | /v1/ask endpoint, plan-bound | | Tenancy model | Multi-tenant cluster or self-host | Shared on Free and Starter; single-tenant on dedicated configurations | | Hosting model | Self-host or managed cloud | Managed-only; shared on Free and Starter, dedicated infrastructure on dedicated configurations | | Operations footprint | One service to operate (or self-run) | One service that replaces row-store + vector + FTS + graph | 06 · operations ## Two different operational stories. Qdrant gives you choice. Self-host the open-source distribution on your own metal, run it in a container, or pay for the managed cloud. The Rust footprint is small enough that running it yourself is a real option for many teams, and the documentation around clustering and snapshots is honest about the trade-offs. For teams that want the source under their control or a self-host requirement to satisfy, that flexibility matters. OriginChainDB is managed-only by design. A dedicated configuration gives each customer a single-tenant database in a region of their choice, with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. The trade-off is real: you get fewer knobs to turn and a smaller ecosystem, but you also do not need a DBA on call to add vector search. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and Supabase Canonical source: https://originchaindb.com/vs/supabase Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and Supabase Compare OriginChainDB and Supabase across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. Supabase has done something genuinely useful for the industry: it took a relational primary engineers already trusted and bundled the boring-but-essential pieces - auth, file storage, realtime channels, edge functions - into one platform. For a relational-first application that occasionally needs an embedding, that bundle is hard to beat. The interesting question is whether the AI surface is a side feature on a relational app, or whether the AI surface is the application. The answer to that question decides the database. choose Supabase if - + You are building a relational-first application where Postgres is the primary store and vectors are a side feature. - + Auth, storage, realtime, and edge functions in one platform are load-bearing for your team's velocity. - + The Postgres ecosystem is non-negotiable - every ORM, every migration tool, every BI connector has to keep working. - + A generous free tier and a first-class dashboard matter for solo developers, prototypes, and side projects. choose OriginChainDB if - + AI features are the workload - embeddings, hybrid search, graph context, natural language are equal citizens to rows. - + You would rather not compose pgvector + ParadeDB + Apache AGE + sync triggers to get one consistent surface. - + Rows, embeddings, full-text postings, and graph edges have to commit atomically in one round-trip. - + You want a managed database - single-tenant on a dedicated configuration, with no shared Postgres pool - and a natural-language endpoint that ships, not bolts on. 02 · where supabase wins ## Postgres with the boring parts already built. Supabase's core insight is that most application teams do not just need a database - they need a database, an auth provider, a file store, a realtime channel, and a place to run a tiny piece of server-side logic. Stitching those together used to mean four vendors and a weekend. Supabase ships them as one platform, with Postgres as the source of truth and the rest exposed through APIs that respect row-level security. For a solo developer or small team, that bundle is the fastest path from idea to deployed application that the industry has yet produced. The Postgres-native posture matters. Every ORM (Prisma, Drizzle, SQLAlchemy, ActiveRecord) talks to a Supabase database the same way it talks to RDS or a self-hosted Postgres. Migrations are plain SQL. Extensions install the way they always do. If you outgrow the platform or want to move, the data and schema are portable in a way that is genuinely rare among modern managed databases. The free tier is generous enough that hobbyists and side projects can ship without paying anything, and the Studio interface - table editor, SQL console, log explorer, auth viewer - is one of the best dashboards in the segment. Vectors are a real capability via pgvector, which now supports HNSW indexes and is genuinely usable for retrieval workloads. Full-text via tsvector is respectable. Realtime subscriptions and Edge Functions plug straight into the same row-level-security model, so a feature like "stream new messages to anyone authorised to see this room" is a few lines of SQL and a JavaScript handler. For a relational-first application where the AI features are additive, Supabase is often the right call. 03 · where originchaindb is different ## AI shapes are first-class. Not extensions you wire up. OriginChainDB is built around a single managed engine. Rows, secondary indexes, vector embeddings, HNSW graphs, BM25 full-text postings, and graph edges all live in that engine. The query engine compiles SQL, vector top-k, BM25 search, graph traversal, and natural-language questions to the same plan tree. There are no extensions to install or version-pin because the AI shapes are not extensions - they are first-class. HNSW has two operating points worth naming: the default `high_recall` mode hits recall@10 = 0.96 at 100k vectors with p99 around 109 ms, and a `fast` mode runs p99 around 37 ms at recall@10 ≈ 0.69. For scale, the IVF-PQ index compresses each vector so a large corpus still fits one box; we have not published a measured result at 100M scale. The consequence is atomicity that crosses shapes. A single insert writes the row, every secondary index entry, every forward and reverse edge, the BM25 postings, and the vector embedding in one atomic commit. With Supabase, the same insert in pgvector + tsvector lives inside a Postgres transaction - which is excellent - but if you are also writing to ParadeDB indexes, Apache AGE edges, or any out-of-process derived state, you are back to writing your own consistency story. OriginChainDB folds the entire derived state into the one commit. On a dedicated configuration, each customer gets a single-tenant database in a region of their choice - its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Supabase pools many projects onto shared Postgres infrastructure, which is the right trade-off for the price point and the workloads they target; OriginChainDB's dedicated configurations make the opposite trade-off, while its Free and Starter plans share infrastructure too. If compute and storage that no other customer shares are part of your compliance or performance budget, that matters. Natural language is part of the same surface. `/v1/ask` compiles an English question to the same plan AST as a hand-written query - same cost model, same EXPLAIN output, same per-node statistics. The model emits a plan; the executor runs it. There is no LLM on the hot path, no token-priced query layer to budget for, and no second service to deploy alongside the database. 04 · the atomicity gap ## What "one INSERT, one atomic commit" actually buys you. The standard architecture for an AI feature on Supabase is a Postgres transaction that touches a row, a tsvector column, and a pgvector column. Inside one transaction, that is genuinely atomic - Postgres's MVCC keeps the three writes consistent. The seam shows up when the AI surface grows. Add an Apache AGE edge for graph context, a ParadeDB index for richer full-text, a search-service mirror for typo-tolerant retrieval, or an out-of-process embedding worker, and the consistency story splits across services. Most teams paper over the gap with retry queues and reconciliation jobs. OriginChainDB folds the entire derived state into the write path. The row, the embedding, the full-text postings, every edge update - all of them are part of the same atomic write, which commits as one unit. Partial writes are dropped on recovery, so there is no half-written state to clean up. That property is verified at runtime: a fault-injection harness deliberately crashes the database at multiple boundaries, and recovery is asserted to equal a prefix of the op stream every time. We run it for a million deterministic iterations on every build. Reads compose the same way. A query can filter on structured columns, rank by vector similarity, intersect with a BM25 search, and join across a graph edge - in one round trip, against one consistent snapshot. With Supabase, a query that combines pgvector + tsvector + a recursive CTE for graph context is possible inside one Postgres statement; once a fourth shape (or a non-Postgres engine) enters the picture, you are stitching results in application code. OriginChainDB exists because that boundary is exactly where AI applications keep getting bitten. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | Supabase | OriginChainDB | | --- | --- | --- | | Core engine | Postgres + extensions | Single managed engine, multi-shape | | Tenancy model | Shared Postgres pool by default | Shared on Free and Starter; single-tenant on dedicated configurations | | Vector search | pgvector extension | Native HNSW + f32 SIMD | | Full-text | tsvector / GIN built-in | Native BM25 + phrase + stemming | | Graph traversal | Recursive CTE | Native fwd / rev edges + Dijkstra | | Atomicity across shapes | Per-row within a Postgres transaction | Row + index + embedding + posting + edge in ONE atomic commit | | Natural-language query | Bring-your-own LLM layer | /v1/ask endpoint, plan-bound | | Auth + storage + realtime | First-class, in-platform | Out of scope - bring your own | | Studio / dashboard DX | Polished, table-editor + SQL console | Admin console + REST + thin SDK | | Replication | Postgres streaming replication | Async standby + fenced automatic failover | | Pricing shape | Free tier + pro / team / enterprise | Free, Starter, or a dedicated compute configuration + flat add-ons | | Operations footprint | Managed Postgres + adjacent services | One service that replaces row-store + vector + FTS + graph | 06 · operations ## Two different operational stories. Supabase is operationally a delight if you have not run a database before. The dashboard shows you exactly what is in your tables, the SQL console executes against the live primary, the auth viewer lets you tweak RLS without leaving the browser, and Edge Functions deploy from `git push`. For a small team, that is hours of plumbing they do not have to write. The trade-off is that the underlying Postgres is shared with other projects on lower tiers, the connection-pool model has well-known sharp edges with serverless callers, and tuning is constrained by what the platform exposes. OriginChainDB is managed-only by design. A dedicated configuration gives each customer a single-tenant database in a region of their choice, with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. The trade-off is real: you give up the breadth of Supabase's bundle (auth, storage, realtime, edge runtime) in exchange for a database where vector / full-text / graph / NL are first-class, and single-tenant on dedicated configurations. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot) --- # OriginChainDB and Weaviate Canonical source: https://originchaindb.com/vs/weaviate Sitemap last modified: 2026-09-23T08:33:23.000Z compare # OriginChainDB and Weaviate Compare OriginChainDB and Weaviate across data models, query capabilities, deployment and workload fit. 01 · choose the right one ## The honest split. Pick the one that matches your workload. Weaviate and OriginChainDB overlap in the parts of an AI application where embeddings meet text - both can serve hybrid retrieval, both expose vectors and BM25 against the same record, both let you filter by structured metadata. They diverge in scope and shape. Weaviate is a vector-and-hybrid search engine with a class schema and a GraphQL surface; OriginChainDB is a full multi-shape database with SQL, vector, full-text, graph, and natural-language all compiling to the same plan tree. The right answer depends on whether your workload is shaped like search-with-metadata or like an application database that also embeds. choose Weaviate if - + You want an open-source vector database you can read, fork, and self-host on your own infrastructure. - + GraphQL is a fit for your application surface - you like the schema-first, single-endpoint shape it gives you. - + Built-in embedding modules (text2vec-openai, text2vec-cohere, multi2vec-clip) are valuable so you do not write embedding code yourself. - + Your retrieval is dominated by vector similarity and hybrid search, with metadata filters that fit on the class itself. choose OriginChainDB if - + Relational queries - JOIN, GROUP BY, HAVING - are first-class alongside vector search, not a secondary concern. - + Rows, embeddings, full-text postings, and graph edges have to commit consistently in one round-trip. - + You prefer dedicated single-tenant infrastructure (a dedicated configuration) to a shared multi-tenant cluster, with no operator burden on your side. - + You want a managed database - single-tenant on a dedicated configuration - that ships natural-language query as a first-class endpoint. 02 · where weaviate wins ## Open-source heritage and a GraphQL surface that fits. Weaviate has been one of the load-bearing pieces of the open-source AI stack for years. The codebase is in the open, the licensing terms make self-hosting a real option, and the community has grown a healthy ecosystem of clients, examples, and integrations. If "we read the source" is on your shortlist of requirements, that is a hard requirement to meet, and Weaviate meets it cleanly. The GraphQL API is a genuine strength when it fits. Class schemas with typed properties, references between classes, and a single endpoint that returns precisely the shape you ask for is a clean fit for many application surfaces - particularly frontends that already speak GraphQL. The hybrid search syntax - vector similarity blended with BM25 by an alpha parameter, all in one query - is one of the more polished hybrid stories in the vector-database space. The module system is also worth naming. Weaviate ships first-class integrations with embedding providers - text2vec-openai, text2vec-cohere, text2vec-huggingface, multi2vec-clip - so you can write a class definition that pulls embeddings on insert without ever shipping the vector from your application. For teams that would rather not maintain embedding code, this is a meaningful productivity win and the kind of opinionated convenience that takes years to design well. 03 · where originchaindb is different ## SQL is a first-class shape. Not a sidecar. OriginChainDB compiles SQL - JOIN across up to 32 tables, GROUP BY, HAVING, OUTER, LIMIT, ORDER BY - to the same plan tree that runs vector top-k, BM25 search, graph traversal, and natural-language questions. There is no second engine for relational work, no sidecar database holding "the rest of the application," and no application code joining results across stores. A query that wants "documents matching this filter, ranked by vector similarity to this question, joined to the author table, grouped by team" is one statement against one managed engine. The HNSW index has two operating points worth naming concretely: the default `high_recall` mode hits recall@10 = 0.96 at 100k vectors with p99 around 109 ms, and a `fast` mode trades recall for latency, running p99 around 37 ms at recall@10 ≈ 0.69 on the same dataset. For scale, the IVF-PQ index compresses each vector so a large corpus still fits one box; we have not published a measured result at 100M scale. You pick the operating point per workload, and the cost model picks between full scan and index scan from query statistics, so SIMD predicates can prune work before vector distance is even computed. Graph is a first-class shape too, not a cross-reference between classes. Forward and reverse edges live in the same managed engine, with native shortest-path (Dijkstra) and traversal primitives. The same query that ranks by vector similarity can hop along an edge - "from this document, walk to its author, find their other recent documents, restrict by topic, rank by similarity to this question" - without leaving the database. Natural language is part of the same surface. `/v1/ask` compiles an English question to the same plan AST as a hand-written query - same cost model, same EXPLAIN output, same per-node statistics. The model emits a plan; the executor runs it. There is no LLM in the hot path, no token-priced query layer, and no second service to deploy alongside the database. 04 · the atomicity gap ## Multi-shape writes in one atomic commit. With Weaviate you usually have a relational primary somewhere - Postgres, MySQL, or another OLTP store - that owns the canonical row. Weaviate then owns the vector and the BM25 posting for that row. Insert a document, embed it, write it to Weaviate, hope nothing crashed in between. Most teams paper over the gap with idempotency keys, retry queues, and a reconciliation job that scans for orphaned rows or orphaned vectors. It works most of the time. OriginChainDB folds the entire derived state into the write path. The row, the embedding, every secondary index, the BM25 postings, every forward and reverse edge - all of them are part of the same atomic write, which commits and replicates to the standby as one unit. Partial writes are dropped on recovery, so there is no half-written state to clean up. That property is verified at runtime: a fault-injection harness deliberately crashes the database at multiple boundaries, asserting recovered state equals a prefix of the op stream every time. For applications where a delete on the row has to also delete its embedding, where a row update has to invalidate a stale vector, or where retrieval has to combine vector similarity with a row-level filter and a graph hop, one managed engine is the cleaner answer. For applications where Weaviate is the only store of record - vectors with metadata, no separate primary - the dual-store consistency problem does not apply, and Weaviate's atomicity within a class is sufficient. 05 · side by side ## The detailed comparison. A capability-by-capability look, so the trade-off is explicit before you pick. | Capability | Weaviate | OriginChainDB | | --- | --- | --- | | Primary use case | Vector + hybrid search with class schemas | Multi-shape DB - rows + vectors + FTS + graph + NL | | API surface | GraphQL + REST | REST + SQL + a thin SDK | | Relational queries | Cross-references between classes | Full JOIN, GROUP BY, OUTER, HAVING, LIMIT | | Embedding modules | Built-in (text2vec-openai, cohere, etc) | Bring your own embedder; store and index here | | Vector index | HNSW (configurable) | HNSW + IVF + IVF-PQ (64× memory) + binary quant (32× memory) + PQ + sparse | | Full-text search | BM25 + hybrid alpha-blend | Native BM25 + phrase + 18-lang stemming + 9-lang lemmas + ICU, atomic with rows | | Graph traversal | Cross-references; not graph traversal | Native fwd / rev edges + PageRank + LPA + Node2Vec + GraphSAGE + Cypher v3 | | Atomicity across shapes | Per-class write semantics | Row + index + embedding + posting + edge in ONE atomic commit | | Tenancy model | Multi-tenant cluster (logical tenants) | Shared on Free and Starter; single-tenant on dedicated configurations | | Natural-language query | External - your LLM layer | /v1/ask endpoint, plan-bound | | Hosting model | Self-host or managed cloud | Managed-only; shared on Free and Starter, dedicated infrastructure on dedicated configurations | | Operations footprint | One service to operate (or self-run) | One service that replaces row-store + vector + FTS + graph | 06 · operations ## Two different operational stories. Weaviate gives you choice. Self-host on your own infrastructure with the open-source distribution, run it on Kubernetes with the official operator, or pay for the managed cloud offering. Each option has its own trade-offs around upgrades, replication topology, and observability, and the choice is yours to make. For teams that have an opinion about where their database runs and want the source code under their control, that flexibility is real. OriginChainDB is managed-only by design. A dedicated configuration gives each customer a single-tenant database in a region of their choice, with its own HTTPS endpoint, its own bearer token, and compute and database storage that no other customer shares. Free and Starter run on shared, scale-to-zero infrastructure instead. We provision, patch, back up, replicate, and upgrade. You post requests, get JSON back. The trade-off is real: you get fewer knobs and a smaller ecosystem, but you also do not need an operator or a DBA on call to add vector search. Failover is structural. A write is acknowledged only after it is flushed to durable storage on the primary, so an acknowledged write is on disk before the response is sent, not buffered in memory. Replication to the standby is asynchronous - the primary streams committed frames continuously but does not block on the standby - so an abrupt loss of the primary can lose the most recent acknowledged writes. How much is at risk depends on how far behind the standby had fallen. Promotion is fenced by a single-primary claim, so split-brain cannot happen; automatic promotion is available but off by default, and it refuses to promote a standby that is not fully caught up. A new standby bootstraps from a consistent snapshot without stalling writes. [For developers Check supported query features](https://originchaindb.com/docs) [For your team Plan a workload evaluation](https://originchaindb.com/pilot)