The State of Apache Arrow in 2026: Ten Years In, the Invisible Standard Is Everywhere
A comprehensive look at Apache Arrow's impact ten years in, covering its origins, technical details, adoption, and future in AI workloads.
A comprehensive look at Apache Arrow's impact ten years in, covering its origins, technical details, adoption, and future in AI workloads.
Explores semantic layers in data platforms, their role in ensuring metric trustworthiness, and their critical importance for AI agent accuracy.
How to use AI coding agents like Claude Code and Codex for deterministic data engineering, ensuring reproducibility and trust.
Explores using PyIceberg without Apache Spark for Python-based Iceberg table operations, focusing on architecture, boundaries, and production patterns.
Explores the real-time lakehouse architecture, focusing on streaming, Iceberg snapshots, and the dual clocks of event arrival and query visibility.
A guide to Apache Iceberg View Federation for portable SQL views across multiple query engines in a data lakehouse.
Overview of modern Python tools for Apache Iceberg, including PyIceberg, IceFrame, and CLI for metadata management.
Explains how Apache Iceberg V3 improves CDC pipelines with deletion vectors and row lineage, solving delete file accumulation.
A practical walkthrough of working with Apache Iceberg on Dremio Cloud, covering table creation, data ingestion, optimization, and AI-powered analytics.
Explains the modular Apache Lakehouse architecture using open-source components like Parquet, Iceberg, Polaris, and Arrow for vendor-neutral data management.
A monthly roundup of interesting links focused on Kafka, event streaming, stream processing, and analytics in the tech industry.
Testing Claude Code's ability to build a production-ready dbt project for a data pipeline, evaluating prompts and skills.
Explores the current capabilities and limitations of using Claude Code (AI) to build a dbt project, arguing it won't replace data engineers yet.
A technical demonstration of using Claude Code AI to autonomously debug and adapt dbt data models by analyzing data anomalies.
A guide to integrating Dremio's data platform with JetBrains AI Assistant for enhanced data querying, pipeline generation, and app development within JetBrains IDEs.
Monthly job board for database professionals, featuring remote and onsite data engineering, DBA, and analytics roles from March 2026.
A monthly roundup of tech links focusing on data engineering, Kafka, AI, and software development, including personal articles and industry news.
A guide to the core principles and systems thinking required for data engineering, beyond just learning specific tools.
A guide to designing reliable, fault-tolerant data pipelines with architectural principles like idempotency, observability, and DAG-based workflows.
Argues that data quality must be enforced at the pipeline's ingestion point, not patched in dashboards, to ensure consistent, reliable data.