IngestThis
BLOG
COMMUNITY
PODCAST

Tag: data engineering

2026-07-25 β€’ Alex Merced

The Apache Iceberg Market in the Middle of 2026

A survey of the Apache Iceberg market in July 2026: the state of the specification, platform support, the acquisition wa...

2026-07-25 β€’ Alex Merced

Three Vendors Are Rebuilding the Path From Transaction to Agent

Databricks, Snowflake, and SAP are closing the gap between operational databases and analytical platforms through acquis...

2026-07-06 β€’ Alex Merced

Deterministic Data Engineering With AI Harnesses: Using Claude Code, Codex, Antigravity, and OpenCode for Data Work You Can Actually Trust

How to use AI agent harnesses for data engineering without losing determinism, reproducibility, and trust in your data p...

2026-07-06 β€’ Alex Merced

The State of Apache Iceberg v4 in July 2026: What the Dev List Tells Us About the Format's Next Chapter

What the Iceberg v4 dev list tells us about adaptive metadata trees, single-file commits, column updates, and the format...

2026-07-06 β€’ Alex Merced

The State of Agentic AI Standards in 2026: MCP, A2A, WebMCP, OSI, and the Protocol Stack Taking Shape

The agentic AI protocol stack is solidifying in 2026 β€” MCP for tools, A2A for agents, WebMCP for the web, OSI for semant...

2026-07-06 β€’ Alex Merced

The State of Apache Arrow in 2026: Ten Years In, the Invisible Standard Is Everywhere

Apache Arrow at 10 β€” ADBC, Flight SQL, nanoarrow, the AI reinterpretation, and how an in-memory standard eliminated the ...

2026-07-06 β€’ Alex Merced

The State of Apache Parquet in 2026: The Quiet Format Enters Its Loudest Decade

Apache Parquet in 2026 β€” variant types, geospatial, ALP encoding, footer redesign, the versioning debate, and how the de...

2026-07-06 β€’ Alex Merced

The State of Apache Polaris in July 2026: From Incubating Catalog to the Governance Layer of the Open Lakehouse

Apache Polaris as a TLP β€” federation, credential vending, semantic layers, lineage, and how the open catalog became the ...

2026-07-06 β€’ Alex Merced

The State of Streaming to Apache Iceberg in July 2026: Every Path, Its Latency, and What to Do When Seconds Are Not Fast Enough

Every path for streaming data into Iceberg in 2026 β€” Flink, Spark, Kafka Connect, broker-native, managed pipelines β€” wit...

2026-05-23 β€’ Alex Merced

An In-Depth Overview of the Apache Iceberg 1.11.0 Release

Apache Iceberg 1.11.0 delivers manifest list encryption, the new pluggable File Format API, credential lifecycle refresh...

2026-05-23 β€’ Alex Merced

Single-Node Data Engineering: DuckDB, DataFusion, Polars, and LakeSail

Optimize single-node data engineering with DuckDB, DataFusion, Polars, and LakeSail. Compare architectures and learn whe...

2026-04-29 β€’ Alex Merced

What Are Table Formats and Why Were They Needed?

Table formats like Apache Iceberg solved the ACID, schema, and performance problems that turned data lakes into data swa...

2026-04-29 β€’ Alex Merced

The Metadata Structure of Modern Table Formats

Iceberg uses a metadata tree, Delta Lake uses a transaction log, Hudi uses a timeline. Here is exactly how each format o...

2026-04-29 β€’ Alex Merced

Performance and Apache Iceberg's Metadata

Iceberg's three-layer metadata tree eliminates directory listing and enables multi-level data skipping. Here is how scan...

2026-04-29 β€’ Alex Merced

Partition Evolution: Change Your Partitioning Without Rewriting Data

Iceberg lets you change partition schemes without rewriting data. Here is how partition evolution works internally and w...

2026-04-29 β€’ Alex Merced

Hidden Partitioning: How Iceberg Eliminates Accidental Full Table Scans

Iceberg's hidden partitioning separates physical layout from user queries using transform functions. Here is how it work...

2026-04-29 β€’ Alex Merced

Writing to an Apache Iceberg Table: How Commits and ACID Actually Work

Here is exactly how an engine writes to an Iceberg table, step by step, from data files through the atomic commit that m...

2026-04-29 β€’ Alex Merced

What Are Lakehouse Catalogs? The Role of Catalogs in Apache Iceberg

Lakehouse catalogs store metadata pointers, manage namespaces, and enforce access control. Here is the complete catalog ...

2026-04-29 β€’ Alex Merced

When Catalogs Are Embedded in Storage

S3 Tables and MinIO AI Stor embed the Iceberg catalog directly in the storage layer. Here is when embedded catalogs make...

2026-04-29 β€’ Alex Merced

How Data Lake Table Storage Degrades Over Time

Iceberg tables degrade through small files, orphan files, metadata bloat, sort order decay, and partition skew. Here is ...

Categories

data engineering
oltp
database
data
frontend
data lakehouse
Data Engineering
Data Lakehouse
Javascript
Data Architecture
Data Analytics
Devops
Data Modeling
DevOps
python
sql
rust
AI
Apache Iceberg
Software Development
Semantic Layer
Agentic Analytics
Agentic Lakehouse
AI & Machine Learning
AI Tools & Software Development
Artificial Intelligence
AI & Agents
Data Platforms
Open Source
Lakehouse
Agentic AI
Apache Arrow
Apache Parquet
Apache Polaris
AI & Society
MCP
Security & Governance
Technology & Culture
Hardware
AI & Security
TopicsData EngineeringApache IcebergData LakehouseAI & Machine Learning
SiteAll ArticlesRSS FeedSitemap
AuthorAlex MercedLinkedInTwitter / X

The Alex Merced Network

  • AlexMerced.com
  • WhoIsAlexMerced.com
  • AlexMercedMedia.com
  • Books
  • AlexMercedCoder.dev
  • AlexMercedData.com
  • DataLakehouseHub.com
  • IcebergLakehouse.com
  • AgenticLakehouse.com
  • SemanticLakehouse.com
  • DataEngnr.com
  • AlexMerced.blog
  • GrokOverflow.com

Events & Community

  • Agentic Lakehouse Events
  • Data Lakehouse Hub Events
  • Data Lakehouse Hub Slack
  • Data Events Slack
  • Data & Tech Slack
  • r/datalakehouseandai
  • Data Lakehouse Hub on LinkedIn
  • Alex Merced Tech
  • Alex Merced Data & AI

Β© 2026 Alex Merced β€” alexmercedcoder.dev

The views, thoughts, and opinions expressed on this site belong solely to Alex Merced and do not represent the views of any organization or employer.