Staff Data Engineer · Distributed Systems · Open Source
I build the data plane behind real time decisions.
I design streaming platforms, query systems, and cloud data infrastructure that remain observable, recoverable, and predictably fast under production load.
Stream processing
Correctness designed into every event path.
Event time, watermarks, checkpoint alignment, backpressure, idempotent sinks, and predictable recovery.
Query systems
Performance that starts with how bytes move.
Columnar formats, vectorized execution, storage layout, memory behavior, and benchmarks grounded in real workloads.
Platform engineering
Infrastructure that operators can understand.
Clear observability, safe deployment patterns, resilient defaults, and cost discipline across cloud data platforms.
A practical stack spanning streaming, compute, storage, orchestration, and observability.
About
Staff level engineering grounded in production behavior, scale, and operational confidence.
My work sits at the intersection of real time systems, cloud data platforms, and practical technical leadership. I care about architecture that performs well in production and stays understandable for the teams operating it.
What I build
Streaming and batch platforms for critical data flows, with a focus on reliability, observability, storage efficiency, and low-latency execution.
How I work
I favor simple, high-leverage architecture decisions, careful tuning, and delivery patterns that help teams move faster without sacrificing confidence in production.
Why it matters
The strongest systems are not just scalable on paper. They recover predictably, stay observable under load, and keep costs under control as usage grows.
Reliable data platforms are built by sweating the runtime details that others skip.
Experience
Experience building real-time and analytics platforms across product and enterprise domains.
The through-line across these roles is consistent: design dependable pipelines, improve performance and cost behavior, and ship systems that downstream teams can operate with confidence.
- Working across enterprise data and platform initiatives with a focus on reliable architecture, clear ownership, and scalable delivery.
- Bringing production discipline from streaming systems into broader data workflows, platform patterns, and technical direction.
- Promoted to lead the Streaming AI data engineering track, guiding architecture and delivery for real-time workloads on AWS.
- Continued optimization of hot-path IO, storage layout, and state handling to improve runtime behavior and infrastructure efficiency.
- Strengthened production readiness through exactly-once sinks, checkpoint strategy, and recovery tuning across Flink-based services.
- Designed real-time processing on AWS using Flink on Kinesis Data Analytics, MSK, DynamoDB, and S3 for high-value production workloads.
- Reduced hot-path IO and delivered major annual cost savings through DynamoDB modeling, payload compaction, and S3 layout tuning.
- Built exactly-once sinks with checkpoint alignment and idempotent upserts while tuning RocksDB state and JVM behavior for recovery and latency.
- Delivered regulated pipelines with secure ingestion, lineage, and data quality gates to improve analytics readiness and operational trust.
- Standardized batch and streaming jobs with reproducible configuration, deployment discipline, and monitoring that reduced delivery friction.
- Built event-driven analytics with Kafka, Spark, and Delta Lake and exposed downstream access through Dremio and REST services.
- Improved query performance with partitioning, Z-ordering, predicate pushdown, and compaction to lower compute and storage cost.
- Integrated Medicare and Medicaid datasets with SQL and distributed data processing to improve revenue capture and reporting readiness.
- Supported analytics workflows with reliable pipeline behavior across Spark, Flink, Kafka, PostgreSQL, and AWS services.
- Delivered optimized ETL on Teradata and Informatica while standardizing SLAs, validations, and delivery quality for healthcare data workflows.
- Built a strong foundation in enterprise data movement, operational rigor, and quality-minded delivery.
Systems
Depth across the complete path from event ingestion to trusted insight.
I work across languages, compute frameworks, storage formats, cloud services, and platform design patterns, with a practical bias toward runtime behavior and production supportability.
01
Languages
Java, Python, Go, Rust, Scala, C++, SQL, and shell used with a practical bias toward maintainability and runtime performance.
02
Streaming and batch
Apache Flink, Kafka and MSK, Spark, Pulsar, Kinesis Data Analytics, and Airflow across event driven and analytical workloads.
03
Storage and infrastructure
Iceberg, Delta Lake, Hudi, Arrow, Parquet, DynamoDB, S3, PostgreSQL, Kubernetes, and Docker for production systems.
04
Query and platform systems
DataFusion, Trino, Presto, ClickHouse, Dremio, DuckDB, observability, and performance with cost optimization.
Architecture focus
- Designing pipelines that stay understandable as they scale in complexity and traffic.
- Keeping throughput, resilience, and cost efficiency aligned instead of trading one against another blindly.
- Making operational behavior visible through stronger monitoring, lineage, and debugging hooks.
Delivery strengths
- Reproducible job configuration, disciplined deployment patterns, and production-minded defaults.
- Hands-on tuning of storage layout, state handling, JVM/runtime behavior, and data models.
- Clear collaboration with downstream analytics, platform, and product teams.
Where I add leverage
- Greenfield streaming architecture and modernization of high-volume legacy pipelines.
- Platform hardening for reliability, observability, and easier incident response.
- Performance optimization efforts that translate directly into lower cloud spend.
Selected work
Small, focused systems built to understand large engineering ideas.
These projects explore columnar execution, stream processing internals, data platform design, systems performance, and developer facing infrastructure.
Systems lab
Dremel
Two compact columnar SQL engines, one in Rust and one in C++, built to answer the same queries over the same bytes.
SpiderOxide
An asynchronous Python crawler framework with native Rust acceleration for performance sensitive work.
DataWizz
Local-first lakehouse and analytics workspace inspired by Databricks, Snowflake, Airflow, and Superset, with file ingestion, SQL exploration, Delta publishing, orchestration, and dashboards.
FlowCore
A Rust stream processing engine with event time, windows, watermarks, late event handling, checkpoints, and a live dashboard.
GoXStream
A Flink inspired stream processor in Go with operator graphs, checkpoints, and connectors for Kafka, files, and databases.
Astra Sentinel
A native Rust desktop workstation for local malware triage, combining multi-hash inspection, YARA scanning, recursive analysis, and JSON reporting.
Education
Jawaharlal Nehru Technological University, Hyderabad
B.Tech in Electrical Engineering
GPA 4.0/4.0 (2014 - 2018)
Professional summary
Staff Data Engineer with hands on depth in streaming systems, cloud data platforms, distributed runtime tuning, and production observability.
Contact
- Emailrohankumardubey497@gmail.com
- GitHubgithub.com/rohankumardubey
- LinkedInrohan-kumar-dubey-3a9a31156
- Portfoliorohankumardubey.github.io
Open source
Learning in public, one system at a time.
Open source is a major part of how I learn and contribute. I read production implementations, reproduce ideas in focused projects, report issues, and share improvements that can help the wider engineering community.
Curiosity becomes more useful when the work is visible.
My GitHub is a living systems notebook covering data engines, stream processors, storage, databases, Rust infrastructure, and experiments inspired by the projects I study.
Start a conversation
Let's make difficult data systems boring to operate.
If you are working on streaming platforms, query engines, open source infrastructure, or large scale data systems, I would be glad to connect.