Stream processing is the discipline of continuous data: real-time data streaming pipelines that ingest, transform, and emit event streams as they arrive rather than on a schedule. Practitioners run Apache Kafka estates, write Apache Flink jobs, and design the event-driven architecture that connects producers to consumers. Demand is being pulled by AI rather than by the analytics team. IBM Think's coverage of Confluent's 2026 Data Streaming Report, drawn from 4,625 IT leaders across 14 countries, found 72 percent saying insufficient infrastructure for real-time data processing stalls their AI scaling, up from 61 percent a year earlier, while only 32 percent have agentic AI in production .
Challenges in Stream Processing Recruiting
Real-time data streaming became the constraint AI programs hit first
The report's numbers describe a discipline being asked to carry infrastructure it was never funded to build. Sixty-six percent of IT leaders cite data lineage, timeliness, and quality uncertainty as an AI-scaling obstacle, 65 percent point at fragmented data ownership, and 71 percent report a skills and expertise gap, up from 66 percent in 2025 . Confluent's own summary adds the investment side: 88 percent of IT leaders rank data streaming as a high investment priority, and 94 percent say streaming increases, or is expected to increase, the impact of their AI investments . Streaming was once a specialist purchase for fraud and personalization. It is now the pipe that agentic systems, vector search, and operational ML all assume exists, and the funding is arriving faster than the people who can operate it.
The operational consequence lands on hiring. A team that runs one Kafka cluster suddenly needs producer standards, consumer contracts, schema governance, and incident response for a system the whole company now depends on. Briefs written before that shift ask for a Kafka engineer. The work on the ground is platform operation under load, and the title pool for those two things is not the same pool.
Apache Kafka topics and partitions are two different jobs
Kafka's own introduction draws the boundary. Each topic is a partitioned log, an ordered, immutable sequence of records with offsets; consumers form groups, and each record is delivered to one consumer instance per group, with ordering guaranteed only within a partition, not across a topic . Partitioning is how the log scales, and keyed messages route to a consistent partition while keyless messages spread round-robin .
That split maps to two careers. Broker-side engineers own replication, retention, storage, and cluster operations. Producer-and-consumer-side engineers own key design, partition counts, consumer groups, offset commits, and rebalance behavior. A resume that says Kafka covers both, and most candidates have genuinely only touched one. The interview that asks which partitions a message went to and why a consumer lagged is the one that finds out.
Apache Flink state and checkpointing separate stream processing owners
Flink's fault-tolerance documentation is the discipline's core text. The engine periodically snapshots every operator's state to durable storage; on failure it rewinds and replays the source streams, and exactly-once means every event affects managed state exactly once, not that every event is processed once . End-to-end exactly-once requires more: sources must be replayable, and sinks must be transactional or idempotent .
These are the sentences that divide the candidate pool. Someone who has run Flink in production can explain checkpoint intervals, state size, barrier alignment, and why a sink had to be idempotent. Someone who has only done the tutorial can explain the SQL. The difference is not depth of knowledge; it is which failures the person has personally survived, and the CV almost never says.
Event-driven architecture pushes contracts into the schema
In an event-driven architecture, the topic schema is the API. Every producer change is a compatibility decision for every consumer, and teams that have not centralized those decisions learn the lesson through outages: a field renamed, a type widened, a consumer that fails silently on an unknown event. The senior profile in this craft is the engineer who treats schemas as versioned contracts and knows which changes are safe without a consumer migration.
Screening for this is cheap and specific. Ask what happens to the downstream when a producer adds a field, removes one, or changes a type, and who owns the compatibility rule. Candidates who answer with the registry and the deployment order have operated a real estate. Candidates who answer with what Kafka permits by default have only read about one.
Late-arriving data punishes low-latency data ingestion pipelines
Low-latency data ingestion is easy to claim and hard to prove, because correctness in streaming is a function of time. Watermarks tell the engine how late an event may arrive; set them tight and you drop legitimate records, set them loose and every window waits on the slowest producer. Late-arriving data, duplicate events, and replayed history all meet in the same code path, and the engineer who has not designed that path discovers it in production.
This is where batch experience fails silently. A nightly job can rerun; a stream cannot rewind a customer's order without duplicating a charge. The screening question is therefore not about throughput. It is about what the pipeline does with an event that arrives two hours after its window closed, and the answer reveals whether the candidate has ever owned a real-time system or only benchmarked one. Ask it twice: once for the happy path, once for the week the producer went down, because the second answer is where low-latency data ingestion claims either hold or dissolve.
Event streams claims collapse under watermark and checkpoint questions
Assessment for this craft runs backward from failure. Take the candidate's largest event streams claim and ask how the pipeline survived: what did the watermark do under a producer stall, what was checkpointed and how large was the state, what happened on a consumer rebalance . Then ask about the money event: an order that was delivered twice, or dropped, and how it was found. Owners narrate these with partition offsets and checkpoint intervals. Tourists narrate the architecture diagram.
The cost of a weak screen shows up as quiet damage. Non-idempotent consumers double-charge customers. A mis-tuned watermark drops orders that only the reconciliation team notices. Consumer lag grows until a downstream SLA breaks, and the seniors who should have screened spend their weeks on forensics instead. The reverse error is just as real: the engineer who carried a Kafka estate through a rebalance storm is often invisible behind a CV that lists the same tools as everyone else. Reading stream processing evidence means hearing the incident history, and that is an engineering judgment, not a keyword match.
References
- Nearly Half of AI Projects Are Stalling Due to Data Problems, New Study Finds — IBM Think. (accessed 2026-09-28)
- Introduction: Apache Kafka Documentation — Apache Kafka. (accessed 2026-09-28)
- 2026 Data Streaming Report: Bridging the Gap Between Data and AI Value — Confluent. (accessed 2026-09-28)
- Intro to Kafka Partitions: Apache Kafka 101 — Confluent Developer. (accessed 2026-09-28)
- Fault Tolerance: Apache Flink Documentation — Apache Flink. (accessed 2026-09-28)
