Splunk, Vega, and Query take three different approaches to federated search. Here is how each one works, what it commits you to, and who pays.
I get a version of the same question every week. How is Query’s federated search different from Splunk’s? And lately: how is it different from what Vega is doing?
The short answer is that the three of us are not building the same kind of product. Splunk federates cloud storage so you can look back at data you chose not to ingest. Vega builds a search index next to your data so you don’t have to move it. Query federates the query itself, across the security stack you already run, and gives you an index only where you decide you want one.
The longer answer is the rest of this post. Every product I’m about to describe builds an index somewhere. What separates them is where that index lives, what it commits you to, and who pays for it. We released Accelerated Connectors recently (see release announcement here and product page here), which means Query now has two answers instead of one, so this is a good time to lay out how all of it works.
Two Lineages
Splunk is a search engine.
Its index is the tsidx file: a term dictionary and posting lists pointing into a compressed journal of raw events, with bloom filters on each bucket to skip the ones that can’t contain a term. Give Splunk a rare string and it will find that string across enormous volumes of data, fast. That capability built a category, and it’s still very good at the thing it was designed for.
Vega is also a search engine. Read their engineering blog and they’ll walk you from a naive word-to-logs map, through tries, to finite state transducers, naming Apache Lucene and tantivy as the implementations that use them. Elsewhere they describe bloom filters, cuckoo filters, and HyperLogLog++ sketches for pruning shards. Different generation, different substrate, same family: an inverted index (they call it a reverse index) with skip structures over it.
Query is a federated query engine: a planner, a cost model, a set of connectors, and predicate pushdown. The lineage is Trino and Calcite, not Lucene.
If you take one thing from this post, take this: search engines need indexes, query engines need connectors.
What an Index Costs You
An index is never free, and the four costs are structural.
It has to be built. That means a delay before your history is searchable, and a pipeline that keeps running forever.
It has to be stored. Whether it’s in a vendor’s SaaS repository or distributed across your own S3 buckets, somebody pays for it every month.
It decides in advance what will be cheap to ask. Tokenization happens at index time, which means somebody guessed at your future questions. Ask something the tokenizer didn’t anticipate and you’re scanning.
It’s bad at analytics. If you don’t believe me, just look at Splunk’s own history.
Splunk rose on inverted indexes. Then customers wanted to aggregate — count by user, group by host, threshold over a window — and everyone discovered that retrieving events and re-extracting fields at search time is a terrible way to compute a GROUP BY.
Splunk’s answer was not to adopt a columnar engine. It was to build more inverted indexes. Data model acceleration produces additional .tsidx files holding pre-extracted, CIM-normalized field-value pairs, and Splunk calls the result the high-performance analytics store. tstats reads those files and never touches the raw journal. That’s the whole trick: the speedup comes from avoiding your raw data, not from a better data structure.
It works. It also costs you a declared data model, CIM mapping as a prerequisite, summary storage next to every bucket, a build period before any of it helps, and a permanent constraint: ask a question outside the model and you’re back to raw search.
Vega made the same architectural bet with a cloud pattern that wasn’t available when Splunk started. Their index is object-storage native instead of assuming local NVMe, and it lives in your own S3, Azure Blob, or GCS bucket in the same region as your telemetry. You aren’t shipping data across a regional boundary to a vendor, and you keep control of storage tiering and residency. Those are real improvements.
Moving an index is not the same as escaping one. It still has to be built, so there’s still a delay before your history is searchable. It still has to be stored, so there’s still a storage line, on your cloud invoice instead of a vendor’s. Tokenization still decides in advance what’s cheap. And it’s still an inverted index, so analytics still need help: Vega has added HLL++ sketches for approximate cardinality and continuously refreshed lookup tables for joins, which are the two things a column store gives you natively.
Where Indexes Win
Rare-indicator hunts. One unusual string, an enormous corpus, and you need to know whether it appears anywhere. An inverted index is the correct data structure for that question. It’s what made Splunk great in the early days, and it’s what Vega is good at now.
Fast hunts over cold data. A hunt across twelve months of unindexed object storage is a scan. An index turns the scan into a lookup. If you run that hunt constantly, the index pays for itself.
Neither of those is in dispute. What’s in dispute is how much of your workload has that shape, and what you give up to be ready for it.
Retention Is a Placement Decision
Say you have a twelve-month retention mandate. A search engine gives you one answer: build the index over all twelve months and pay for it every month, whether you query it once a day or once a quarter. Splunk’s answer is ingest. Vega’s answer is an index in your bucket. Same commitment, different address.
A query engine gives you a different kind of answer: put each slice of data where its cost matches its mission, and search all of it the same way.
Park the archive in whatever S3 tier you like and connect it. Nothing to build, nothing to store beyond the data you were already keeping, no monthly line for the privilege of being able to search it. You pay when you run a query and nothing when you don’t. For an archive you touch a few times a quarter, which is what archive retention usually looks like in practice, that beats twelve months of index storage to speed up four queries.
If some of that data needs to be fast, move that slice somewhere fast. If you’re optimizing for rare-indicator hunts, put it in OpenSearch. If you’re optimizing for analytics, put it in Databricks or ClickHouse. If it’s the twenty sources your scheduled detections hit every five minutes, put it in an Accelerated Connector, which is our own column store, priced on what you keep. Every one of those is a connector, the query language doesn’t change, and the planner decides which one to ask.
That’s the flexibility security architects keep asking us for. The retention mandate is fixed. Where the data sits to satisfy it, and what each tier costs, should be yours to decide, and you should be able to change your mind later without re-platforming.
No Index at All
Query’s original federated search delegates the query to the platform already holding your data. There’s no index, so there’s nothing to build, nothing to store, and no waiting. A source is searchable the moment you connect it, across whatever history it retains, including archive tiers.
The engineering lives in the planner. It pushes down as much of the query as each source can execute and asks for the minimum needed to answer you.
You might be saying: sure, but federated search is slow, and it’s limited to whatever the source can do. I hear this constantly. Both halves are wrong.
We’re as fast as the source, and you choose the source. Federated search over a cold object store is a scan. Federated search over OpenSearch is an index lookup. Federated search over Databricks is a columnar aggregate. Query reaches 62 integrations today, and the speed of any given question is a placement decision you make, not a property of our product you’re stuck with.
Results stream, and no source blocks another. Union queries, which is most of security search, return results as each source produces them. Your time to first result is set by your fastest source, not your slowest. Global sort uses a streaming merge, so the first row lands once every source has produced one; it never waits for a source to finish. Joins and exact global aggregates do have to block, and they’re the exception.
Substring, regex, and CIDR work everywhere. Every field, plus CIDR matching on every IP field. Where a source can run the predicate, we push it down. Where it can’t, we keep the predicate and evaluate it ourselves, paging the source until your limit is genuinely satisfied instead of handing back whatever survived one arbitrary page.
The model is always correct, sometimes slower. That’s the right trade for hunting, because hunting is by definition the thing nobody anticipated. The alternative is fast for what the tokenizer expected and awkward otherwise, and to their credit, Vega’s own engineering blog says so directly: neither Lucene nor tantivy supports efficient suffix or infix search on the same inverted index, and the workaround they describe is maintaining extra indexes over reversed words or word rotations.
And we reach things an index can’t. An index-based architecture can only index what it can write an index next to, which means object stores and systems that already have a query engine of their own. An enormous amount of security-relevant data lives in a system of record with nothing but an API: threat intel platforms, identity providers, CMDBs, EDR consoles, ticketing systems, SaaS audit logs. We treat every one of those as a first-class source, with the same query language and the same predicates as an S3 bucket.
That last point is Query’s breadth advantage. Splunk Federated Search reaches data stores — object stores, lakehouses, and warehouses. Vega indexes cloud object storage and delegates to legacy SIEMs. Query federates the stack: object stores, SIEMs, lakehouses, EDR, identity, cloud control planes, ticketing, and threat intel, in one query.
What Splunk Federates
Since the question I’m asked most is the Splunk one, here’s the direct comparison, built entirely from Splunk’s own documentation.
Splunk’s federated search has grown into a family. Splunk-to-Splunk lets one deployment search another, authenticating to the remote as a configured service account. Beyond that, Splunk Cloud can now federate to Amazon S3, Azure Blob and ADLS Gen 2, Azure Databricks, Snowflake, your own DDSS archive buckets, and Amazon Security Lake. That’s real expansion and it’s worth crediting: a year ago the honest list was S3.
Here’s what the documentation asks of you.
- An AWS-resident Splunk Cloud. Every non-Splunk source requires Splunk Cloud running in an AWS region, including the Azure ones. Your data can live in Azure; your Splunk can’t. Splunk Enterprise doesn’t get any of it, on-prem or BYOL, and neither do FedRAMP High or DoD IL5 stacks — so the most regulated customers are excluded today from the feature designed to save them money.
- A catalog in the middle. Each dataset needs a catalog entry before Splunk can search it. The original S3 experience required an AWS Glue table per dataset; that path is deprecated as of Splunk Cloud 10.4.2604, and as of 10.5.2605 you can’t create new S3 federated indexes at all — existing ones are read-only and have been migrated into the Data Management app. The replacement infers schema and manages the catalog for you, which is a genuine improvement. It’s still a layer that has to be right before a query works.
- A different query language. The federated sources run SPL2, not the SPL your analysts know. Federated Analytics searches use a third command, sdselect.
- A separate meter. Data Scan Units, 10 TB of scan each, purchased annually on top of the Splunk Cloud subscription, one pool shared across the federated sources. Amazon Athena scans the same S3 for about $5 per terabyte.
- A lookback lane. Splunk positions this for infrequently searched data. Real-time detection and investigation stay on ingested data, and Federated Analytics caps its data lake indexes at 31 days, which tells you where the hot path is meant to live.
Storage and warehouses. Not GCS, not your EDR console, not your identity provider, not anything that lives behind an API.
None of that is hidden; it’s how the product is documented. But it matters when you’re asking “what do I stop ingesting,” because the answer they usually need is “most of it,” and a lookback lane over your object stores and warehouses, priced as a scan meter on top of the platform bill, doesn’t get them there. Federating the archive is table stakes. Federating the stack, and running detections in place over data you never ingested, is something an ingest-priced business model has no reason to build.
What You Get by Not Owning the Data
Three of our structural advantages exist for one reason: we don’t hold your data. None of them can be bolted onto an ingest-first product after the fact.
Enrichment is live, and there’s no pipeline. Ask which group a user belongs to and we ask your IdP, right then. Ask who owns an asset and we ask your CMDB. Ask whether an indicator is known and we ask your threat intel platform, which matters most of the three, because freshness is the value of intel. A nightly-refreshed indicator table is worthless against a campaign rotating infrastructure every few hours.
Everybody else solves this by copying your system of record into their world on a schedule. Splunk uses lookups and KV store collections refreshed by scheduled searches. Vega built Organizational Context and Dynamic Lookup Tables, which they describe as continuously refreshed reusable assets. That’s a pipeline. A small, well-managed, vendor-operated pipeline, with a staleness window. An index-based architecture doesn’t have a choice here, because you can’t build an index over an API you don’t control.
One caveat: live enrichment gives you current state. For “is this account still active” or “who owns this host,” current state is exactly what you want. For “who had access at the time of the incident,” you want state as of the event, and that’s a case where a snapshot serves you better.
Predicates are uniform, for the reason above. No tokenizer had to guess.
Sensitive fields are never requested at all. Our connector ReBAC and column masking are enforced at plan time. A field that’s masked, or that sits in a connector the user can’t access, isn’t fetched and then hidden. It is never asked for. It doesn’t cross the wire, doesn’t enter our memory, doesn’t show up in our logs, and never lands in a result set.
If you’ve ever taken a security questionnaire to your legal team, you know why that matters. “The field never enters the vendor’s system” is a better answer to an auditor than “the vendor’s interface hides it.” A product that already ingested and indexed your data cannot write the first sentence, however good its access controls are.
It also makes your queries cheaper, because fewer columns means fewer bytes scanned at the source. I don’t often get to ship a security control that makes things faster.
How Do You Know Your Search Was Complete?
Plenty of people have written about federated search performance. But do you want to know what nobody writes about, even though everybody should? Whether the answer you’re looking at is the whole answer.
Every acceleration strategy and every federation strategy creates this problem. Splunk’s version is familiar to anyone who has run summariesonly=true against data model summaries that hadn’t caught up: you silently get fewer results, and the number on your screen looks as authoritative as a correct one. Vega streams partial results, which is genuinely good, but they haven’t published what they tell you about per-source completion. That’s the dangerous case for a hunter. You can’t tell “not present” from “hasn’t reported yet,” and a partial negative on an indicator search is the kind of thing that ends up in an incident review.
In Query, when a source warns, errors, throttles, or truncates, we surface it in the UI. When we’re applying a filter the source couldn’t run, we page that source until your limit is actually satisfied rather than stopping at the first page. Ordering and limiting are there so you can probe a result that looks suspicious.
Our Own Index
Federation is bounded by the source. When a source is slow, expensive to query, aggressively rate-limited, or getting hammered by scheduled detections every minute, the right answer is to stop asking it.
Accelerated Connectors build a columnar index: values stored and compressed by column, held in physical sort order, with block-level statistics that let the engine skip data it never needs to read. It’s still an index with its own (fast) ingestion and another bill. But it’s a column store, which is the structure that fits where inverted indexes struggle, and it’s data you opt to index.
Columns have types. Values are integers, timestamps, and IP addresses, not strings in a lexicon. Numeric ranges, arithmetic, sums, averages, and percentiles are native operations.
Every column is there. No data model to declare, no CIM mapping to maintain as a precondition. Physical sort order is a performance decision, not a capability gate. You can’t ask a question that falls outside the model, because there is no model.
Vectorized execution and per-type compression, neither of which a file format designed for term lookup can give you.
Splunk’s answer to analytics was a tsidx of pre-extracted fields. Vega’s is sketches and lookup tables layered on an inverted index. Ours is a column store.
The Bill
Splunk built this market on ingest pricing, and that one decision created the market’s central problem. When you’re billed on volume, your only cost lever is data reduction, and data reduction is a standing headcount commitment. Every Splunk shop I’ve talked to spends a lot of energy deciding what they’ll stop collecting. Splunk has since added workload-based pricing, and the treadmill it created is still running in most of the accounts we walk into.
One data point: a Fortune 500 customer cut its SIEM ingestion from 10 TB a day to 4 TB a day after moving the rest to federated access, which they value at about $2M in SIEM cost over three years.
Query Accelerated Connectors are priced on storage. Per terabyte, per day, for what you chose to keep. It’s predictable, and it moves in the same direction as a number you already track. Queries against accelerated data are free at the margin, so your hunters can iterate all afternoon without anybody getting a phone call about the bill.
Federated search has no storage cost. You pay whatever the target platform charges to run the query, which is honest, and which makes it the cheaper answer below a certain query volume and the right answer for anything you touch rarely.
Vega hasn’t published what they meter. That’s a fair thing to ask any vendor right now, when the entire industry is pointing autonomous agents at its security data. If an agent investigates every alert, and every investigation is a hunt across cold storage, somebody should be able to tell you what a thousand of those a day costs before you sign. Ask them. We’ll tell you ours.
Two Modes
The workload mix inside a single enterprise spans the whole range. Ad-hoc investigation. Scheduled detections firing every five minutes. Twelve-month archive hunting. Enrichment against a dozen systems of record. No single point on the curve is right for all of it.
Splunk arrived at the same conclusion, which I think is the strongest evidence available. Federated Analytics for Security Lake keeps a short-retention hot index for high-frequency detection while older data stays put and gets reached by ad-hoc federated search. Hot tier plus federated archive. Two modes.
So having two modes isn’t the differentiator. The differentiators are that ours share one language, one set of semantics, and one planner deciding which to use; that our federated mode reaches the whole stack, not just a list of data stores; and that our second mode is a column store instead of another inverted index.
Credit Where It’s Due
The inverted index has had a remarkable run. Lucene shipped in 1999, Splunk built a category on the same idea a few years later, and for most of this century nearly every serious log search product has been a variation on a term dictionary and a posting list. It earned its place by working, over and over, at scales nobody anticipated when it was designed.
Vega gave it a modern home. Moving the index into object storage instead of local disk fixes something that needed fixing for a lot of customers. A better home doesn’t change what the structure can reach.
It still won’t beat not building an index at all, when what you need is cheap retention rather than fast hunting. And it still won’t query your IdP, your CMDB, or your threat intel platform at the moment you ask. You can’t build an index over somebody else’s API.
Whatever you’re evaluating, ask: where does the index live, what does it commit you to, and who pays? If you want more choice and control over your own answer, let’s connect. There’s a good chance we can help.
