Sponsored by

What’s in today’s newsletter:

Amazon Aurora PostgreSQL queries Iceberg tables efficiently 🚀

AWS integrates DuckDB to speed PostgreSQL queries ⚡

Cloudflare Basin enables serverless, fast data analysis ☁️

Top 6 Data Lineage Platforms for Regulated Sectors⚠️

Xmentium boosts Microsoft Fabric with AI document intelligence 🤖

Also, check out the weekly Deep Dive - Database use in AI

What would you delegate if you had a whole team?

With Skydive, you can build a team of AI agents to take work off your plate.

Create agents for customer support, sales, marketing, engineering, ops, and more. Give each one a role, connect the tools they need, and hand off the work.

Your agents can work on their own or together, sharing context and handing off tasks to get bigger jobs done.

Start with one agent. Build a whole team around you.

AWS

TL;DR: Amazon Aurora PostgreSQL now supports querying Apache Iceberg tables in S3 using Parquet, enabling seamless, consistent OLTP and OLAP workloads without data movement, enhancing analytics efficiency and interoperability.

  • Amazon Aurora PostgreSQL now supports querying Apache Iceberg tables stored in Amazon S3 using Parquet format.

  • This integration enables SQL queries on large analytic datasets without moving or transforming data.

  • Users benefit from transactional consistency and combined OLTP and OLAP workloads in a single system.

  • The feature reduces data duplication and latency, enhancing analytics efficiency with open standards interoperability.

Why this matters: This integration unifies transactional and analytical workloads in Aurora PostgreSQL, enabling direct SQL access to large data lake tables without data movement. It reduces complexity, duplication, and latency, accelerating insights and fostering flexible, interoperable data architectures that support modern enterprise analytics and decision-making.

TL;DR: AWS integrated DuckDB into PostgreSQL to speed complex queries on live and historical data, simplifying architecture and boosting analytics performance within a single system for faster, seamless insights.

  • AWS has embedded DuckDB into its PostgreSQL DBMS to accelerate queries on live and historical data.

  • DuckDB's in-process OLAP capabilities allow faster complex SQL analytics without moving data between systems.

  • The integration simplifies database architecture and improves performance for mixed operational and archival queries.

  • This move enhances AWS’s analytics offerings and may inspire wider adoption of embedded analytical databases.

Why this matters: AWS’s integration of DuckDB into PostgreSQL significantly boosts query speed and simplifies architecture by enabling seamless analysis of live and historical data. This advancement improves real-time decision-making, reduces infrastructure complexity, and may set a precedent for future high-performance, embedded analytical database solutions.

DATA PLATFORMS

TL;DR: Cloudflare Basin is a serverless data analysis tool enabling developers to run SQL on large data with no infrastructure, integrating with Cloudflare Workers for low-latency, real-time, scalable edge computing.

  • Cloudflare Basin is a serverless data analysis tool that eliminates infrastructure management for developers.

  • Developers can write SQL queries on data formats like Parquet, leveraging scalable compute power on demand.

  • Basin integrates with Cloudflare Workers, enabling low-latency applications with edge function and real-time log analysis.

  • The tool improves development speed, cost efficiency, and scalability, enhancing decision-making with fast localized processing.

Why this matters: Cloudflare Basin democratizes data analysis by removing infrastructure challenges and enabling scalable, real-time querying with low latency. This accelerates development, lowers costs, and enhances decision-making, especially for edge applications, marking a significant advancement in developer tools and serverless data solutions.

Work With Cloud Database Insider

Looking to reach CTOs, CIOs, and enterprise Data Engineers and Data Architects?

Limited sponsorship slots available each month.

DATA LINEAGE

TL;DR: Six top data lineage platforms in 2026 help regulated industries ensure data accuracy, compliance, and traceability with automation, visualization, integration, and security, enhancing audit readiness, risk management, and decision-making.

  • Data lineage platforms are crucial for regulated industries to ensure data accuracy, compliance, and traceability in 2026.

  • Six top platforms offer automated lineage extraction, visualization, integration, and strong security for complex data systems.

  • Supporting standards like GDPR and HIPAA, these platforms help industries rapidly respond to audits and compliance requirements.

  • Adoption improves risk mitigation, data governance, analytics, and decision-making, boosting operational integrity and stakeholder trust.

Why this matters: As regulations tighten, adopting advanced data lineage platforms is vital for regulated industries to ensure compliance, prevent costly errors, and protect reputation. These tools enhance transparency, security, and data quality, enabling organizations to respond swiftly to audits and make informed decisions, maintaining competitive advantage and operational integrity.

MICROSOFT FABRIC

TL;DR: Microsoft Fabric now integrates AI-driven Copilot, autonomous Agents, and app-building tools to simplify data workflows, boost efficiency, and empower both technical and non-technical users in enterprise analytics.

  • Microsoft Fabric now integrates enhanced AI-driven Copilot features for intuitive, code-free data interaction and insights generation.

  • AI-powered Agents in Fabric autonomously handle complex data tasks, improving efficiency and streamlining analytics operations.

  • New app-building tools in Fabric empower users to create AI-embedded custom applications for business decision support.

  • These updates embed AI deeply into Fabric, simplifying data workflows and accelerating innovation in enterprise analytics.

Why this matters: Microsoft Fabric’s integration of AI-driven Copilot, autonomous Agents, and app-building tools fundamentally lowers barriers to data analysis and application development. This accelerates enterprise innovation, enhances efficiency, and democratizes access to AI-powered insights, enabling businesses to make smarter, faster, and more collaborative data-driven decisions.

TL;DR: ClickHouse now integrates natively with Microsoft Fabric’s OneLake, boosting data access and analytics performance while highlighting Microsoft’s open ecosystem strategy and enhancing enterprise data workflow efficiency.

  • ClickHouse has integrated natively with Microsoft Fabric's OneLake, enabling seamless data access without duplication.

  • The collaboration improves operational efficiency and enhances analytics performance on OneLake data for enterprises.

  • Microsoft Fabric users can now leverage ClickHouse’s high-performance analytics engine in their data workflows.

  • The partnership highlights Microsoft’s commitment to an open ecosystem and unifying diverse analytics tools.

Why this matters: The integration of ClickHouse with Microsoft Fabric’s OneLake eliminates data silos, boosting efficiency and accelerating real-time analytics. This partnership further cements Microsoft’s open ecosystem approach, enabling enterprises to combine best-in-class open-source and cloud technologies for more agile, insightful data-driven decision-making.

TL;DR: Xmentium partnered with Microsoft Fabric to integrate its Document Intelligence platform, automating unstructured data extraction and enabling AI-ready analytics, enhancing operational efficiency and accelerating digital transformation.

  • Xmentium’s Document Intelligence automates extraction from complex documents, reducing manual processing time significantly.

  • As an IQ Sharing launch partner, Xmentium integrates document data seamlessly into Microsoft Fabric’s analytics workflows.

  • The partnership enhances accessibility of unstructured data, enabling smarter decisions using AI and machine learning in Fabric.

  • This collaboration advances digital transformation, improving efficiency and insights with integrated document intelligence in Microsoft Fabric.

Why this matters: Integrating Xmentium’s AI-powered document intelligence with Microsoft Fabric transforms unstructured data into actionable insights, accelerating digital transformation. This collaboration boosts operational efficiency, enhances analytics capabilities, and empowers businesses to leverage comprehensive data workflows, showcasing Microsoft’s commitment to versatile, AI-driven enterprise solutions.

EVERYTHING ELSE IN CLOUD DATABASES

DEEP DIVE

Database use in AI

If you have been reading this newsletter for any length of time, you would notice a skew towards the expanded role of databases far beyond OLTP and classic Data Warehouses.

I am seeing “on the ground” uses in my day to day for Machine Learning, and expanded use in Agentic AI. The use case that is always top of mind for me is the use of databases for the statefulness of agentic AI activity.

I think that the statefulness use case is something I am always thinking about is because of my DBA history that I have mentioned so many times. Just the idea of an agent spinning up a database at will and destroying it at will, without the classic governance that that we are bound to in the enterprise.

The following is a bit of research that covers what I have mentioned above.

The two workloads side by side

The two uses differ on almost every axis an architect cares about, which is why they are converging on different engines.

Dimension

Use 1: Model training

Use 2: Agentic state

Position relative to the model

Upstream: data feeds the model

Downstream: the model drives the data

Primary client

Pipelines, data engineers, training jobs

The agent itself, at machine speed

Dominant access pattern

Large sequential scans, batch writes, append-mostly

Small point reads and writes, read-modify-write loops

Optimised for

Throughput and cost per TB

Latency, isolation and transactional correctness

Data volume per store

TB to PB in a few large stores

MB to GB in many small stores

Mutability

Immutable snapshots, versioned tables

Constantly mutating, frequently branched or rolled back

Lifespan

Years; datasets retained for reproducibility

Seconds to months; many stores are disposable

Typical engines

Lakehouse table formats (Delta, Iceberg), object storage, warehouses, feature stores

Serverless Postgres, SQLite/embedded stores, Redis-class caches, vector indexes

Who provisions

Platform teams, via change control

Increasingly the agent or the agent framework

Core governance question

What data trained this model, and was it permitted?

What did this agent read, change, and on whose authority?

Failure mode

Silent data quality or lineage errors baked into weights

Corrupted state, runaway sprawl, over-privileged actions

Use 1: Databases for model training

In training, the database is the governed, versioned system of record for what a model learned from. It is rarely in the GPU hot path; its value is curation, reproducibility and provenance.

The training data lifecycle

  1. Ingestion. Operational systems, event streams, documents and external corpora land in object storage or a lakehouse. CDC from OLTP systems is the usual enterprise feed.

  2. Curation and quality. Deduplication, filtering, PII scrubbing, toxicity screening and quality scoring run as SQL or Spark jobs over tables. Quality scores become columns that later drive dataset selection.

  3. Labelling and annotation. Labels, human review outcomes and preference pairs (for RLHF or DPO-style post-training) are stored as structured tables keyed to source records.

  4. Feature engineering (classical ML). Feature stores split into an offline store for training and an online store for low-latency inference. Point-in-time joins prevent future data leaking into training rows.

  5. Dataset assembly and versioning. A training set is a pinned snapshot: table time travel, Iceberg branches and tags, or tools such as lakeFS and Nessie. The dataset version is recorded against the model version.

  6. Consumption by the training job. Data loaders stream columnar or sharded files (Parquet, Lance, WebDataset-style shards) straight from object storage. The query engine is out of the loop at this point.

  7. Experiment and model metadata. Runs, parameters, metrics and model registry entries live in a relational backend (MLflow's tracking store is the common example).

  8. Evaluation and feedback. Evaluation sets, benchmark results and production feedback are stored as tables and fed into the next training cycle.

What "database" really means here

For foundation-model pretraining, the store is object storage plus open table formats plus a catalog. The database role is metadata, lineage and access control rather than serving bytes. For enterprise classical ML and fine-tuning, the warehouse or lakehouse is literally the training source, queried directly.

Properties that matter most

  • Reproducibility: immutable snapshots so a model can be traced to the exact data it saw.

  • Point-in-time correctness: features as they were at prediction time, not as they are now.

  • Provenance and consent: licensing, retention and opt-out status carried at row or file level, because removing data after training is expensive.

  • Throughput and cost per TB: separation of storage and compute, columnar formats, cheap object storage.

  • Multimodal support: images, audio, video and embeddings stored alongside structured attributes, increasingly in the same table.

Market signal

Platform vendors are pulling the feature layer inward. Databricks acquired the real-time feature store Tecton in 2025 and positioned it explicitly for serving context to agents, not only to classical models. That framing is itself evidence that the training stack and the agent stack are starting to share infrastructure.

Use 2: Databases for agentic AI

For agents, the database is the memory, the checkpoint and the transaction log. A model call is stateless; everything an agent carries across steps, sessions or failures has to live in external state.

Why agents need state

  • Statelessness of the model: each inference call starts from whatever context is passed in. Continuity is an application concern, not a model feature.

  • Finite, expensive context: the full history cannot be replayed into every call, so state must be stored, summarised and selectively retrieved.

  • Long-running work: agent tasks span minutes to days, cross deployments and survive crashes only if progress is persisted.

  • Interruption by design: human approvals, tool timeouts and retries pause runs mid-flight. Resuming requires a durable checkpoint.

  • Side effects: agents change real systems. Those changes need transactional guarantees and an audit trail.

The state tiers

Agent state is not one thing. Production systems separate it into tiers with very different lifespans and storage needs.

Tier

What it holds

Typical lifespan

Typical store

Key requirement

Working context

Current task variables, tool outputs, scratch results

Seconds to minutes

In-process memory, Redis-class cache, embedded SQLite

Latency

Session history

Messages and turns in one conversation or thread

Hours to days

Postgres, key-value store

Ordered append and read

Execution checkpoints

Graph or workflow state after each step

Run lifetime plus retention

Postgres, Redis, SQLite

Durability; resume and replay

Semantic memory

Facts, preferences, entity knowledge across sessions

Months or longer

Postgres with vectors, vector DB, graph DB

Retrieval quality; update and dedupe

Episodic memory

Records of past runs and their outcomes

Months

Event tables

Queryable history

Procedural memory

Learned instructions, skills, prompt versions

Long-lived, versioned

Versioned tables or files

Version control

Task workspace

Data the agent builds or modifies for a task

Task lifetime

Ephemeral database or branch

Isolation; disposability

Audit trail

Every read, write, tool call and decision

Compliance retention

Append-only tables, traces in the lakehouse

Immutability

Frameworks now encode this split directly. LangGraph's persistence model separates thread-scoped checkpointers (conversation continuity, human-in-the-loop, time travel, fault tolerance) from cross-thread stores for long-term memory, with Postgres as the usual production backend.

Checkpointing and time travel

Checkpointing turns an agent run into a replayable sequence of states. An operator can inspect any prior step, fork from it, or resume after a failure without redoing completed work. This is the agent-era equivalent of a transaction log plus point-in-time recovery, applied to reasoning rather than rows.

Multi-agent coordination

When several agents share a task, the database becomes the coordination layer: shared task state, queues, leases and idempotency keys. Databricks reported that multi-agent system usage on its platform grew 327% in four months (Forbes, Feb 2026). Concurrency control matters more here than in single-agent chat, because agents act in parallel on the same state.

Agents acting on systems of record

Agents increasingly write to operational data through tools, so classic OLTP concerns return: ACID guarantees, row-level security and least privilege. Identity is the new piece. Snowflake made a dedicated SERVICE_AGENT user type generally available in July 2026, giving agents their own identity and privileges, and followed with Restricted Session Scope, a privilege ceiling that caps what an agent can do on a user's behalf.

Agent state is becoming a managed product

In September 2026 Databricks put managed agent memory and managed agent sessions into Beta. Both are Lakebase-backed stores usable from any agent framework: one for durable long-term memory with semantic search, one for conversation history an agent reads and appends to each turn. The pattern is clear: agent state is moving from hand-rolled tables to a platform service.

Why Postgres keeps winning this tier

One Postgres engine can hold transactional state, JSON documents, vectors and full-text indexes, which covers most tiers above in one system. Models also generate SQL and Postgres DDL fluently, which lowers the friction for agents that provision and query their own stores. Both Databricks and Snowflake now ship managed Postgres inside their platforms, which the next section covers.

Ephemeral databases: spin up, use, destroy

Agents now create most databases on the platforms built for them. On Neon, agent-created databases rose from 0.1% in October 2023 to 80% by October 2025, and agents created 97% of database branches (SaaStr summary of the Databricks State of AI Agents report). Databricks co-founder Reynold Xin summarised the shift: humans don't create and destroy thousands of databases a day, but agents do (Forbes).

Why agents create and discard databases

  • App generation. Each app an agent builds gets its own database. When Create.xyz launched its developer agent on Neon, it created 20,000 databases in 36 hours (Madrona).

  • Parallel hypothesis testing. An agent working a hard problem may spin up dozens of isolated environments, test competing approaches, keep the winner and tear down the rest within seconds.

  • Safe work against production-shaped data. Branch a production snapshot, run a migration or a destructive test, inspect the result, discard the branch.

  • Rollback and retry. If an agent's logic fails, it can rewind the database to the moment before the failure and retry in a fresh branch.

  • Isolation between agents. Giving each agent its own copy avoids contention when many agents hit the same data at once (InfoWorld).

  • Portable agent state. Turso's AgentFS puts an agent's entire runtime (files, key-value state, tool-call history) in a single SQLite file, with an in-memory mode for fully ephemeral runs. Snapshotting the agent means copying the file.

What made it economically possible

  • Separation of storage and compute. Stateless compute heads attach to shared storage, so a new database does not mean a new server or a data copy.

  • Copy-on-write branching. A branch shares pages with its parent until it diverges, so a full-fidelity copy of a large database appears in seconds and costs only the changed pages.

  • Scale-to-zero and usage pricing. An idle database costs roughly its storage. Databricks framed the goal as keeping the cost of thousands of ephemeral databases proportional to the queries they run (press release).

  • API-first provisioning. Create, branch, restore and delete are programmatic calls an agent can make, not tickets a DBA processes.

  • Embedded engines. SQLite-class databases run in-process, so the lightest ephemeral state never touches a server at all.

The ephemeral lifecycle

An ephemeral database follows a short, branching lifecycle: provision, work in isolation, evaluate, then either promote the result or destroy the branch.

Only the promoted result survives; the branches and their compute disappear, so the audit record has to be written before teardown.

Vendor landscape (as of October 2026)

Product

Architecture

Status

Agent-relevant capabilities

Databricks Lakebase (built on Neon)

Serverless Postgres, compute separated from lake storage

GA on AWS Feb 2026; GA on Azure Mar 2026

Branching, autoscaling, scale-to-zero, instant restore, Unity Catalog governance; backs managed agent memory and sessions

Snowflake Postgres (from Crunchy Data)

Managed community Postgres, dedicated VM per instance

GA Feb 24, 2026

Runs inside the Snowflake security boundary; positioned for agents over transactional data with Cortex AI

Embedded SQLite-compatible, one file per agent

Beta, open source

In-memory ephemeral mode, tool-call audit in SQL, snapshot by file copy

Distributed SQL with vector search

GA

Vendor reports 90% of new daily clusters are created by agents

The acquisitions behind the first two rows set the market's price for this capability: roughly $1 billion for Neon and $250 million for Crunchy Data (SaaStr). Note the architectural split: Lakebase is built for many short-lived, branchable databases, while Snowflake Postgres's dedicated-VM model favours steady, isolated workloads over high-churn ephemerality.

Read the numbers with care

  • Vendor telemetry. The headline figures come from platforms with a commercial stake in the trend. Neon's former CEO described the early data as coming from a very specific segment of the market (Forbes).

  • Created is not retained. A share of databases created counts throwaway branches; it says little about where durable enterprise data lives.

  • Enterprise lag. The same Databricks report found only 19% of organisations had deployed agents at all (SaaStr). The 80% figure describes the leading edge, not the average estate.

The connective tissue: inference-time retrieval

Retrieval sits between your two uses: data prepared like training data, but served at agent latency. It gives a model fresh, permissioned knowledge without changing its weights.

Why it belongs to both worlds

  • Built like Use 1. Chunking, embedding, indexing and refresh are batch pipelines over governed tables, with the same lineage and quality concerns as training data.

  • Served like Use 2. Each query is a low-latency point lookup, filtered by the caller's permissions at request time, inside an agent's loop.

  • Feature stores are the structured equivalent. An online feature store serves structured context to a model the way a vector index serves documents, which is why Databricks positioned Tecton for agents.

Convergence into general-purpose engines

Vector search is becoming a column type and an index, not a separate database. Postgres covers it through extensions such as pgvector. Databricks ships Lakebase Search (hybrid vector and full-text inside Postgres) and, as of September 2026, its ai_search function can retrieve directly from Lakebase synced tables, alone or alongside vector indexes. Snowflake made multi-index Cortex Search generally available in March 2026 as a tool for its agents.

The loop back to training

Agent state does not stay in Use 2. Traces, tool calls, human corrections and outcomes are exactly the material needed for evaluation sets and fine-tuning data. Platforms are already landing that telemetry in governed tables: Databricks Apps telemetry, generally available in September 2026, persists traces, logs and metrics to Unity Catalog tables. Today's agent audit trail is tomorrow's training corpus, so it needs training-grade provenance from the start.

Architecture implications and risks

The architect's job shifts from curating a few platforms to setting policy for a fleet of databases that no human provisions by hand. Most of the new risk sits in Use 2.

Risk

What goes wrong

Controls

Database sprawl

Thousands of orphaned databases and branches; nobody knows which copies hold what

Time-to-live by default; mandatory owner, agent and task tags; automated reaping; per-agent quotas

Leakage through branches

A branch of production copies sensitive data into a test-grade environment

Branch from masked snapshots; inherit catalog policies on every branch; block branching of restricted tables

Over-privileged agents

An agent inherits a human's broad role and acts beyond its task

Dedicated agent identities; privilege ceilings below the user's role; read-only by default; narrow tool scopes

Cost runaway

Agent loops create, scale or query without bound

Scale-to-zero; per-agent budgets; query tags for cost attribution

Memory poisoning

Wrong or injected facts persist in long-term memory and resurface in later runs

Provenance on every memory write; expiry and review; separate trusted and untrusted memory

Lost evidence

An ephemeral database is destroyed and the record of what happened goes with it

Write audit trails to a durable append-only store outside the ephemeral database before teardown

Erasure gaps

Personal data scattered across sessions, checkpoints, memory, branches and traces

Map every state tier in the retention policy; erasure jobs that reach all of them

Drift between operational and analytical copies

Synced tables diverge from the system of record

Declare one system of record per entity; prefer managed sync over hand-built pipelines

Design principles

  1. Ephemeral by default, durable by promotion. Anything an agent creates expires unless a policy or a person promotes it.

  2. Keep the state tiers distinct. Scratch state and session history should never quietly become a system of record.

  3. Make agent identity first-class. Every read and write should be attributable to a specific agent, acting for a specific principal, on a specific task.

  4. Put the audit trail outside the blast radius. The log of what an agent did must outlive the database it did it in.

  5. Govern branches like their parents. One catalog and one policy set should cover a database and every copy of it.

  6. Design for the loop back to training. Capture traces with consent and provenance metadata now, because they will become training and evaluation data.

What to watch over the next 12 to 24 months

The direction is set; the open questions are about who standardises it and how far it spreads beyond developer tooling.

  • Enterprise spread. Whether agent-created databases move from the vibe-coding and developer segment into regulated enterprise estates, given only 19% of organisations had deployed agents at the time of the Databricks report.

  • Hyperscaler response. Whether native cloud Postgres services match branch-per-agent economics, or cede this tier to data-platform vendors.

  • Managed memory APIs. Whether managed agent memory and session stores, such as Databricks' September 2026 Beta, converge on common interfaces or stay platform-specific.

  • Agent identity standards. Snowflake's SERVICE_AGENT type already supports workload identity federation, including SPIFFE/SPIRE. Watch for cross-platform agent identity and delegation models.

  • One substrate or two. Whether a single governed platform ends up serving both training data and agent state, as the lakehouse-plus-lakebase pitch argues, or whether the two stay on separate engines.

  • Per-agent versus per-task. Whether the database-per-agent model holds at enterprise scale, or settles into short-lived branches off a small number of governed parent databases.

  • Audit expectations. How regulators and auditors treat agent actions on systems of record, which will shape retention, lineage and identity requirements for Use 2.

Sources

Gladstone Benjamin

🎯 Level Up Your Data Career: Database Certification Strategy Call

Struggling to figure out which cloud database certifications actually move the needle for your career and salary?

Skip the guesswork. Book a 1-on-1, 45-minute Database Certification Strategy Call directly with Gladstone Benjamin. Drawing from 27+ years as a Data Architect and DBA and holding multi-cloud certifications across AWS, Azure, GCP, OCI, Snowflake, and Databricks, I’ll help you map out a targeted, high-ROI certification roadmap tailored to your specific background and goals.

Special offer for Cloud Database Insider readers: Save 25% OFF your session!

👉 Book Your Certification Strategy Call Here (Use promo code SEPTCERT2026 at checkout)