A Kafka message body is an opaque blob in the value column; most often JSON, sometimes Avro or protobuf. Whatever the format, most of the engineering effort around βstreaming analyticsβ is, in practice, the work of turning that blob into columns somewhere a query engine can use them.
The usual answer is another pipeline: a Flink or Spark job that consumes the topic, parses each message, flattens the fields and writes them to a table. It works, but now you own a consumer group, a schema mapping, a checkpointing strategy and a dead-letter path; all so that ππ·π΄ππ΄ ππππππππ_ππ = πΊπΈ can prune instead of scan.
Topics to Tables is VAST DataBaseβs answer to that. Because the VAST Event Broker immediately stores topics as tables inside the database, the parsing step can happen at the point of ingest: when a producer writes a message, the parsed columns are written in the same database transaction. You describe the columns you want; VAST fills them in as messages land. The result: no streaming engine to run and a table thatβs queryable within minutes of enabling it, And because itβs the VAST DataBase doing the work, it scales with the cluster rather than a separate streaming tier.

The pipeline you eliminate: traditional streaming requires a separate processing pipeline (left), while Blob Expansion parses data at ingest inside the database (right).
To be clear, this is additive - the traditional approach still works. The VAST Event Broker is Kafka compliant, so Flink, Spark, Kafka Connect and the rest still work against VAST exactly as they do today. And because the expanded table is an ordinary VAST DataBase table, Spark, Trino and other engines can query it directly. Blob Expansion simply means you no longer need a separate job for the most common case: turning a message body into columns.
A note on names. In the product and CLI this feature is called blob2table; the database object it creates is a blob expansion. This post uses βBlob Expansionβ throughout. It ships in VAST Cluster 5.5.0 and builds directly on the Event Broker, which turns each Kafka topic into a VAST DataBase table (Kafka metadata plus the message itself). The first release parses JSON; support for further formats is on the roadmap.
What It Does
You define an expansion on a Kafka topic. The definition has three parts:
The source: the topic (and, later, a column in an ordinary table). The source is a Kafka topic and the source column is the message value.
The target table:Β where the expanded columns live. In the initial release this is a separate table in the same bucket as the topic.
The expansion schema: an Arrow schema naming the fields to pull out of the JSON and the type each should have.
From then on, every message produced to the topic is parsed synchronously and the matching fields are written to the target table as typed columns. Filter and projection pushdown work on them exactly as they would on any other VAST DataBase table. The parsing happens inside VAST, so there is nothing to deploy, tune or monitor on your side.
What Ends Up in the Target Table
The target table gets more than just the fields you asked for, and the extras are what make it useful operationally:
Your expanded columns, typed as declared.
The remaining Kafka columns: key, headers, timestamp and so on are copied across so the row is self-describing.
Partition ID and offset, so any row in the target can be traced back to the exact message in the topic.
Optionally, the original message as-is, if you want the raw payload alongside the parsed fields.
Diagnostic columns that record what happened during parsing (more on those below).

Anatomy of an expanded row: your typed columns, the copied Kafka columns, partition and offset for traceability, and the diagnostic counters.
Nested JSON is handled either by writing nested struct columns or, if you set the flatten nested structs option, by promoting nested fields to top-level columns named with a delimiter (πππππππ__ππππ’, for example). Fields inside arrays stay nested.
Bad Data Doesnβt Stop the Stream
Streaming pipelines break most often on the message that doesnβt match the schema. Blob Expansion is built on the assumption that this will happen and that dropping the message is the wrong response.
Every row carries a ππππππ_πππππ flag and a count of type mismatches. Optionally, two further counters record how many fields in the source JSON had no expansion definition (βexcessiveβ) and how many defined fields were absent from the message (βmissingβ). A single row can have both. The message is never lost, it is still in the topic, and the target row exists with whatever could be parsed.
Parse failures also raise an alert and increment metrics, and a predefined analytics report summarizes expansion health per definition: rows expanded cleanly, rows that failed on invalid JSON, rows that failed on a type mismatch, and rows with excess or missing fields.
Schema Evolution
Schemas drift. The columns you declare are created in the target for you when you first define the expansion: you point it at an empty table and from there the definition is something you evolve rather than recreate:
Add columns to an existing expansion; missing columns are created in the target automatically.
Drop columns from the definition. The column stays in the target table, the system never removes data columns on your behalf.
Drop the expansion entirely, again without touching the targetβs columns.
Inspect the definition at any time to see the mapping and which tables are associated.
Some guard rails apply. Target columns that are part of an expansion canβt be renamed, and the diagnostic columns canβt be dropped while the expansion exists. One expansion per source column, and one source per target table.Β The target table records which source feeds it, and the server enforces this.
Governance and Management
Expansion is a first-class VAST DataBase object. It is managed through REST, the GUI, the CLI and Terraform, with the GUI presenting the target layout the same way it presents table column definitions. Four new identity-policy actions cover create, alter, drop and get, so who can wire a topic into a table is controlled with the same policies as the rest of the database. Definition changes and the resulting inserts are audited.
Snapshots capture both the source table and the expansion definition. Under replication, the expansion process only runs where the bucket is writable, at the replication source or a standalone bucket. The data still replicates normally: the destination receives the topic and its expanded table like any other replicated tables, and carries the expansion definition so it's ready if it ever becomes writable; it just doesn't re-parse rows that were already expanded upstream.
The Trade-Off
Parsing JSON at ingest time is not free, and the design is honest about that. Because expansion is synchronous, it reduces the number of messages per second a topic can absorb, and the cost grows roughly linearly with the number of expanded columns. Copying the original message into the target as well costs more than expanding fields alone. There is a latency budget per message, and a cluster-wide switch to disable expansion during ingest if you need to recover headroom.
That is the same compute you would otherwise spend in a separate streaming job, it has simply moved to where the data is, without the second copy or the operational surface area.
Initial Scope
The first release is deliberately narrow:
Source: Kafka topics only, expanding the message value column
Format: JSON only today, with support for more formats (such as Avro) on the roadmap
Target: an existing table in the same bucket as the topic
No renaming or dropping of the source topic or target table while an expansion exists (drop the expansion first)
Coming in 5.5.3. An optional error topic can be configured on an expansion. Once set, any message that fails parsing is also mirrored to that topic. Because every Event Broker topic is itself a table, you get a queryable record of all parse failures in one place, without having to filter the target table for ππππππ_πππππ = ππππ . 5.5.3 also introduces mandatory fields: [one line on behavior, e.g. fields you mark as required, where a message missing them is treated as a parse failure rather than a row with a "missing" count].
The design already accounts for the next steps.Β More source formats such as Avro, expanding a JSON column of an ordinary table into that same table with re-expansion on update, and further source types, but those follow once the Kafka path is solid.
Looking Ahead: What's Next for Blob Expansion
Storage-level features usually make data cheaper to keep. This one makes it cheaper to use. If the reason your streaming data lives in a separate warehouse is that someone had to write a job to flatten it, that job can go. Produce JSON to the topic; query columns from the table.
In Part 2 we do exactly that on a live 5.5.0 cluster: stand up a Kafka-enabled view, point an expansion at an orders topic, produce JSON and query typed columns, then deliberately feed it malformed messages and watch the stream keep running.
