gourses

← Volver a la búsqueda

DP-750: Microsoft Azure Databricks Data Engineer Associate

udemy · IT y software · ⭐ 4.68 (133 reseñas) · All Levels · en · ⏱ 25,5 h

Impartido por Ansh Lamba JSR · 1.201 alumnos

34.99 USD

Entra en tu cuenta para guardar este curso y volver a él cuando quieras.

Comparar este curso
Ver curso en udemy

Enlace de afiliado: podemos cobrar comisión, sin coste extra para ti. Más información

Descripción

Are you ready to become an AZURE DATABRICKS DATA ENGINEER? Whether you're a beginner or a working professional who wants to level up, this course will guide you step by step with a hands-on, practical, and engaging approach. GAIN STRONG HANDS-ON WITH: • Core Components of Azure Databricks - Gain a deep understanding of Azure Databricks Lakehouse architecture and core components, including Delta Lake, Unity Catalog, Metastore, Volumes, Lakehouse Federation, and UDFs/UDTFs. • CI/CD Bundles - Develop and deploy Databricks Asset Bundles (DABs) for CI/CD workflows using Databricks CLI while implementing Git-based version control with Databricks Git Folders and Azure DevOps. • Databricks Compute - Configure Databricks compute for performance, including autoscaling, node sizing, auto-termination, pools, Photon, cluster policies, library management, and access control. • Data Ingestion - Master data ingestion with Lakeflow Connect, Notebooks, and Azure Data Factory from Azure SQL, Data Lake, REST APIs, and Event Hubs. Handle advanced scenarios including CDC, Auto Loader schema evolution, stream-static joins, watermarking, and late-arriving data. • PySpark & SparkSQL - Implement batch and streaming data transformations using PySpark and Spark SQL. Master data cleansing, joins, aggregations, set operations, pivoting, merges, and DBUtils for file management, secrets, and parameterization. • Lakeflow Declarative Pipelines - Build Lakeflow Declarative Pipelines using Streaming Tables and Materialized Views to create Star Schemas, implement Slowly Changing Dimensions (SCD), and enforce data quality with Expectations. • Job Orchestration - Orchestrate and schedule workflows using Databricks Jobs and Azure Data Factory, implementing retries, error handling, task dependencies, triggers, notifications, task values, and job recovery operations. • Databricks SQL Warehouse - Master Databricks SQL Warehouse features, including query parameters, caching, snippets, alerts, AI/BI Genie, CTAS, COPY INTO, deep/shallow cloning, and query history. • Performance Optimization - Optimize Databricks performance and cost with caching, partitioning, Z-ordering, liquid clustering, VACUUM, deletion vectors, and time travel. Analyze and troubleshoot Spark workloads using DAG visualization and Spark UI. • Security and Governance - Implement Databricks security and governance using Unity Catalog, RLS, data masking, ABAC, Azure Key Vault, service principals, lineage, access controls, permissions management, and data retention policies. • Monitoring and Audit Logging - Configure monitoring, audit logging, and performance tracking with Databricks, Azure Monitor, and Log Analytics. Create alerts and dashboards, optimize costs, and enable secure data sharing with Delta Sharing. What Makes This Course Different? • Super Engaging Lectures - No boring theory here! I explain every concept in a clear and beginner-friendly way using real-life examples and doodle visuals. • Deep Dive into Every Topic - I don’t just scratch the surface. You'll understand the “why” and “how” behind every feature. • Strong Hands-On Focus - Learn by doing! From pipelines to notebooks to warehouse, you’ll build real solutions step-by-step, just like an Azure Databricks Data Engineer does. DISCLAIMER : This course is independently created and not affiliated with or endorsed by Azure Databricks. All content, including explanations and practice materials, is original and intended solely for educational purposes. It does not include any actual certification exam questions and is based on publicly available documentation, real-world scenarios, and personal experience. Product names, logos, and trademarks used are the property of their respective owners and are included only for identification and learning. Always refer to official Azure Databricks documentation for the latest and most accurate information.

Lo que aprenderás

  • Understand the Azure Databricks Lakehouse architecture and its components including Delta Lake, Unity Catalog, Metastore, Volumes, Managed and External Tables
  • Build & Deploy Declarative Automation Bundles (DABs) for CI/CD workflows using Databricks CLI, Databricks Git Folders and Azure DevOps
  • Configure different compute types and performance settings including node count, autoscaling, termination, pooling with Photon engine, and cluster policies
  • Master data ingestion with Lakeflow Connect, Notebooks, Azure Data Factory, from various sources including Azure SQL, Data Lake, REST APIs, and Azure Event Hubs
  • Implement data transformation on both batch and streaming data using PySpark and SparkSQL. Handle Duplicates, Nulls, Filter, Joins, Unions, Except, Pivot, Merge
  • Build Lakeflow Declarative Pipelines with Streaming Tables and Materialized Views to create STAR Schema, Slowly Changing Dimensions (SCD Type) with Expectations
  • Orchestrate and schedule jobs and workflows using Databricks Jobs and Azure Data Factory. Implement error handling, retries, repair, restart, stop and alerts
  • Explore Databricks SQL Warehouse and learn how to work with Query Parameters, Query Caching, Query Snippets, SQL Alerts, AI/BI Genie
  • Optimize with caching, partitioning, Z-ordering, Liquid Clustering, VACUUM, Deletion Vector. Tackle Spilling, Skewing, and Shuffle issues using DAG and Spark UI
  • Secure and govern data with RLS, Data Masking, ABAC, Azure Key Vaults, Service Principals. Manage data lineage, access control, users & groups, retention policy
  • Implement monitoring and audit logging with Databricks, Azure Monitor, and Azure Log Analytics. Enable Delta Sharing to manage data sharing with external orgs

Requisitos

  • Basic SQL knowledge will be required
  • Basic Python programming knowledge will be required
  • No AZURE DATABRICKS knowledge is required - Everything is covered from SCRATCH