Building Trusted Data Platforms with Azure Databricks and GenAI - Second Edition: A Hands-On Guide to Creating Governed Data Products in a Lakehouse
暫譯: 使用 Azure Databricks 和 GenAI 建立可信資料平台 - 第二版:湖倉中創建受管資料產品的實作指南

Kukreja, Manoj

  • 出版商: Packt Publishing
  • 出版日期: 2026-08-07
  • 售價: $2,260
  • 貴賓價: 9.5$2,147
  • 語言: 英文
  • 頁數: 774
  • 裝訂: Quality Paper - also called trade paper
  • ISBN: 1806679779
  • ISBN-13: 9781806679775
  • 相關分類: Power BI
  • 尚未上市,無法訂購

相關主題

商品描述

A practical guide to building a modern, GenAI-powered data platform with a Lakehouse foundation, covering MDM, data mesh, AI enablement, streaming pipelines, observability, and cloud-driven architectures for trusted analytics.

Key Features:

- Discover characteristics of future-ready platforms - data mesh, automation, & observability

- Design trustworthy data products with contracts, federated governance, and decentralized ownership

- Understand how GenAI accelerates Lakehouse development and enables self-service analytics

Book Description:

Discover the defining hallmarks of future-ready data platforms, including data mesh architectures, intelligent automation, and end-to-end data observability. Learn how to design and deliver trusted data products through data contracts, federated governance, decentralized domain ownership, and endorsed datasets. The book explores modern Lakehouse patterns with a strong focus on the medallion architecture, explaining how bronze, silver, and gold layers transform raw data into analytics-ready assets governed through Unity Catalog. You'll gain practical guidance on MDM linkages, survivorship rules, and entity resolution to ensure consistent master data across domains. It also covers real-time and streaming pipelines that integrate seamlessly with the Lakehouse. We focus on self-service analytics, showing how governed data products let business users explore, analyze, and derive insights independently with confidence. Finally, understand how GenAI accelerates platform development through automated code generation using tools like Claude Code and Databricks Genie Code, enabling faster pipeline creation, governance, and analytics delivery.

What You Will Learn:

- Future-ready platforms: data mesh, automation, observability

- Design trusted data products with contracts and governance

- Build Lakehouses with medallion architecture: bronze, silver, gold

- Apply Unity Catalog for governance and endorsed datasets

- Implement MDM using linkages, survivorship, and entity resolution

- Develop real-time and streaming pipelines at scale

- Enable governed self-service analytics for business users

- Use GenAI to generate code with Claude and Databricks Genie

Who this book is for:

This book is crafted for aspiring data and AI/ML architects, engineers and analysts starting their data engineering journey and seeking a practical, hands-on guide to building scalable, cloud-driven data platforms. It's ideal for professionals familiar with PySpark who want to design modern Lakehouse architectures using Delta Lake, while learning MDM, data mesh, AI enablement, streaming pipelines, automation, and data observability. A working knowledge of Python, Spark, and SQL is expected.

Table of Contents

- The Story of Data Engineering and Analytics

- Discovering Storage and Compute in Lakehouses

- Data Engineering on Microsoft Azure

- Designing Future Data Platforms

- Databricks, Medallion Architecture & Delta Lake

- Understanding Modern Data Pipelines

- Data Collection Stage - The Bronze Layer

- Data Curation Stage - The Silver Layer

- Data Aggregation Stage - The Gold Layer

- Next-Gen Data Analytics with Generative AI

- Data Observability

- Data Governance

商品描述(中文翻譯)

一個實用的指南,教你如何建立一個現代化的、以 GenAI 為驅動的數據平台,基於 Lakehouse 基礎,涵蓋主數據管理 (MDM)、數據網格、AI 啟用、串流管道、可觀察性以及以雲為驅動的可信分析架構。

主要特點:
- 探索未來準備平台的特徵 - 數據網格、自動化與可觀察性
- 設計可信的數據產品,包含合約、聯邦治理和去中心化擁有權
- 了解 GenAI 如何加速 Lakehouse 的開發並啟用自助式分析

書籍描述:
探索未來準備的數據平台的定義特徵,包括數據網格架構、智能自動化和端到端數據可觀察性。學習如何通過數據合約、聯邦治理、去中心化的領域擁有權和經過認可的數據集來設計和交付可信的數據產品。本書探討現代 Lakehouse 模式,強調獎章架構,解釋如何將銅層、銀層和金層轉換原始數據為可分析的資產,並通過 Unity Catalog 進行治理。你將獲得有關 MDM 連結、生存規則和實體解析的實用指導,以確保跨領域的一致主數據。它還涵蓋與 Lakehouse 無縫集成的實時和串流管道。我們專注於自助式分析,展示如何讓業務用戶在有信心的情況下獨立探索、分析和獲取見解的治理數據產品。最後,了解 GenAI 如何通過使用 Claude Code 和 Databricks Genie Code 等工具自動生成代碼來加速平台開發,從而實現更快的管道創建、治理和分析交付。

你將學到的內容:
- 未來準備的平台:數據網格、自動化、可觀察性
- 設計可信的數據產品,包含合約和治理
- 使用獎章架構建立 Lakehouse:銅層、銀層、金層
- 應用 Unity Catalog 進行治理和經過認可的數據集
- 使用連結、生存和實體解析實施 MDM
- 大規模開發實時和串流管道
- 為業務用戶啟用治理的自助式分析
- 使用 GenAI 生成代碼,搭配 Claude 和 Databricks Genie

本書適合對象:
本書專為有志於成為數據和 AI/ML 架構師、工程師和分析師的人士而設,特別是那些剛開始數據工程之旅並尋求實用、動手指南以建立可擴展的雲驅動數據平台的專業人士。它非常適合熟悉 PySpark 的專業人士,想要使用 Delta Lake 設計現代 Lakehouse 架構,同時學習 MDM、數據網格、AI 啟用、串流管道、自動化和數據可觀察性。預期具備 Python、Spark 和 SQL 的基本知識。

目錄:
- 數據工程與分析的故事
- 探索 Lakehouse 中的存儲和計算
- 在 Microsoft Azure 上的數據工程
- 設計未來的數據平台
- Databricks、獎章架構與 Delta Lake
- 理解現代數據管道
- 數據收集階段 - 銅層
- 數據策展階段 - 銀層
- 數據聚合階段 - 金層
- 下一代數據分析與生成式 AI
- 數據可觀察性
- 數據治理