Preparing Your Data Lakehouse for the EU AI Act: Auditable Lineage and Data Provenance Requirements
Explains how to prepare a data lakehouse for EU AI Act compliance, focusing on data lineage and provenance.
Alex Merced — Developer and technical writer sharing in-depth insights on data engineering, Apache Iceberg, data lakehouse architectures, Python tooling, and modern analytics platforms, with a strong focus on practical, hands-on learning.
610 articles from this blog
Explains how to prepare a data lakehouse for EU AI Act compliance, focusing on data lineage and provenance.
Technical guide on preparing data lakehouses for EU AI Act compliance, focusing on auditable lineage and data provenance.
Analysis of stateless MCP patterns for data platforms, focusing on agentic lakehouse architecture and Dremio integration.
A guide to evaluating agentic analytics tools for enterprise AI, focusing on governance, semantics, and production readiness.
Explains the need for machine-readable metric contracts to standardize business meaning before AI agents access data platforms.
A deep dive comparing block vs. object storage, explaining how lakehouses made slower object storage fast for analytics.
Deep dive into designing custom AI agent harnesses, covering architecture layers like loops, tools, context, and control for production systems.
How to use AI coding agents like Claude Code and Codex for deterministic data engineering, ensuring reproducibility and trust.
Historical comparison between 18th-century suppression of Diderot's Encyclopédie and modern debates over open AI models, exploring knowledge access politics.
Explores event-driven table compaction using agents in lakehouse architectures, focusing on small file problems and production patterns.
Explores how schema contracts, semantic models, and governed lakehouse design improve Microsoft Fabric agentic analytics and AI behavior.
Article on fine-grained security for AI agents, focusing on identity, masking, and policy enforcement in machine-speed analytics.
Analysis of Apache Iceberg v4 performance focusing on metadata round trips, root manifests, and object storage latency for platform engineers.
Analysis of LTAP (Lakehouse Transactional Analytical Processing) focusing on freshness, isolation, and workload boundaries for data architects.
Analysis of how enterprise AI value comes from governed context and data contracts, not just model choice.
Explores using PyIceberg without Apache Spark for Python-based Iceberg table operations, focusing on architecture, boundaries, and production patterns.
Explores the real-time lakehouse architecture, focusing on streaming, Iceberg snapshots, and the dual clocks of event arrival and query visibility.
Analysis of REST Catalog V2 LoadTable and client capability negotiation for lakehouse platforms.
Analysis of Rust vs C++ for building native Iceberg scan operators, focusing on production performance, safety, and interoperability.
Explores server-side commit deconflicting in REST catalogs for high-concurrency lakehouse writes, focusing on architecture, specs, and operational patterns.