Portfolio Assistant
Portfolio Assistant
Projects • Skills • Experience
Gunasekaran Ravi · Data Engineer
OPEN TO DATA ENGINEERING ROLES

GunasekaranRavi

Data Engineer | Azure • Databricks • PySpark • SQL

Data Engineer with 4+ years of experience building Azure-based ETL/ELT and Lakehouse data pipelines using Azure Data Factory, Databricks, PySpark, Delta Lake and SQL.

CHENNAI, INDIA
Azure Data Engineering illustration featuring Azure Data Factory, Databricks, Data Lake Storage and Azure SQLAzure Data Engineering illustration featuring Azure Data Factory, Databricks, Data Lake Storage and Azure SQL
Lakehouse Engineering Lifecycle — Applied Across All Five Projects
Ingest
Bronze raw ingestion
Transform
Silver cleansing & standardization
Validate
Data quality & reconciliation
Model
Gold facts & dimensions
Monitor
Audit logging & pipeline health
Serve
SQL & business analytics
About Me

Engineering Scalable Data Platforms
With PySpark & Delta Lake

I'm Gunasekaran Ravi, a Data Engineer at Opseazy Solutions Private Limited (Aug 2022 — Present), based in Chennai, India. I work on data pipeline development using PySpark and SQL, building Lakehouse-oriented ETL/ELT workflows.

Outside of my day-to-day role, I build and publish independent Databricks Lakehouse projects implementing Medallion Architecture, Delta Lake, Delta MERGE incremental processing, data quality validation, and pipeline monitoring. Each project maps to a target Azure production architecture, with what's actually implemented shown honestly alongside what is planned as reference architecture.

Pipeline Engineer
PySpark + Delta Lake Lakehouse pipeline design
Lakehouse Engineer
Databricks, Delta Lake, Medallion Architecture
Analytics Engineer
SQL analytics, business KPIs, data modeling
Gunasekaran RaviGunasekaran Ravi
Gunasekaran Ravi
Data Engineer
CORE STACK
Azure Data Factory
Azure Databricks
PySpark
SQL
Delta Lake
ADLS Gen2
Technical Skills

Data Engineering Skill Set

Solid badges = core expertise used across flagship projects
Cloud & Azure
PROFESSIONAL SKILL SET
Microsoft AzureAzure Data FactoryAzure Data Lake Storage Gen2Azure Databricks· 4 yrs

See Architecture section for implemented vs. target usage per project.

Data Engineering
PIPELINES & DATA MODELING
ETL/ELTMedallion Architecture· 4 yrsIncremental ProcessingCDCSCD Type 2Data QualityData ReconciliationDimensional Modeling
Processing & Lakehouse
DISTRIBUTED PROCESSING
Apache Spark· 4 yrsPySpark· 4 yrsDelta Lake· 4 yrsSpark Structured StreamingUnity Catalog
Programming & Querying
LANGUAGES & QUERY
PythonSQL· 4 yrsSpark SQL
Analytics
BI & VISUALIZATION
Power BI
Engineering Tools
VERSION CONTROL & CI
GitGitHubGitHub Actions
Engineering Pattern

Lakehouse Architecture

The Medallion pattern is implemented across all five projects — Bronze, Silver, and Gold layers — built with PySpark and Delta Lake, with incremental processing, data quality, and monitoring built in.

STEP 01
Source Data
Generated / Raw Input

Raw source data and generated inputs.

Structured Records
STEP 02
Bronze Layer
Raw Ingestion

Raw ingestion with metadata retained.

Raw IngestionDelta Lake
STEP 03
Silver Layer
Cleansed & Trusted

Cleaned, standardized and validated data.

Data Quality Checks
STEP 04
Gold Layer
Business-Ready Analytics

Business-ready facts, dimensions and KPIs.

Delta MERGE
STEP 05
SQL / Analytics
Consumption Layer

SQL and analytical consumption.

Spark SQL
CROSS-CUTTING CAPABILITIES
Incremental Processing (Delta MERGE / CDC)Data Quality ValidationReconciliationAudit LoggingPipeline MonitoringEnd-to-End Validation
TARGET AZURE ARCHITECTURE

This is the target production Azure mapping for the Databricks, PySpark, and Delta Lake implementation — not a claim that it has been fully deployed. Only the Databricks, PySpark, and Delta Lake layer is implemented and validated with execution evidence in these repositories; Azure Data Factory, ADLS Gen2, and Power BI are shown here as reference architecture only.

Source Systems
Azure Data Factory
ADLS Gen2
Azure Databricks
Delta Lake
SQL / Power BI
Experience

Professional Experience

Data Engineer
Aug 2022 — Present
Opseazy Solutions Private Limited
  • Designed and maintained Azure Data Factory pipelines for batch ingestion using linked services, datasets, triggers and integration runtimes
  • Implemented Copy, Lookup, ForEach, Web and Notebook activities, with incremental loading via watermark columns and last-modified timestamps
  • Developed Azure Databricks notebooks in Python and PySpark for cleansing, joins, aggregations, window functions and business-rule transformations
  • Implemented Delta Lake MERGE/UPSERT, schema evolution and Bronze/Silver/Gold Lakehouse processing with SQL-based data quality and reconciliation checks
  • Automated and monitored Databricks jobs and data pipelines, investigated failures, and maintained technical documentation
Portfolio Projects

Selected Personal Data Engineering Projects

Independently designed and implemented outside of employment five PySpark/Delta Lake Lakehouse pipelines, each with its own GitHub repository and execution evidence, demonstrating Medallion Architecture, incremental processing and data quality patterns.

★ FLAGSHIP
FLAGSHIP PROJECT · COMPLETED
BATCH LAKEHOUSE · SALES ANALYTICS

Enterprise Sales Lakehouse

Production-style Databricks Lakehouse demonstrating Medallion Architecture, Delta MERGE incremental processing, data quality, reconciliation, monitoring and business analytics for enterprise sales data.

Nine-notebook Bronze → Silver → Gold pipeline with Delta MERGE incremental processing and full data-quality, reconciliation, audit and validation coverage.
PySparkDelta LakeDatabricksMedallion ArchitectureDelta MERGEGitHub Actions
Source Generation
Bronze (Raw)
Silver (Trusted)
Gold (Analytics)
SQL / BI Consumption
COMPLETED
LAKEHOUSE · CDC & SCD TYPE 2

Real-Time E-Commerce Lakehouse

Databricks Lakehouse processing e-commerce orders, customers and products through CDC-based incremental upserts and SCD Type 2 dimension history, backed by a full data-quality and audit framework.

An 11-stage Delta MERGE-based CDC incremental pipeline paired with SCD Type 2 historical dimension versioning.
PySparkDelta LakeCDCSCD Type 2DatabricksGitHub Actions
Source Generation
Bronze
Silver
CDC + SCD Type 2
Gold Aggregations
COMPLETED
STRUCTURED STREAMING · IOT TELEMETRY

Real-Time IoT Streaming Lakehouse

Spark Structured Streaming Lakehouse processing IoT sensor telemetry with event-time watermarking, windowed aggregations and anomaly detection across Bronze, Silver and Gold layers.

Event-time watermarking with 5-minute windowed aggregations and multi-dimension anomaly detection.
PySparkSpark Structured StreamingDelta LakeWatermarkingDatabricksGitHub Actions
IoT Source
Bronze Streaming
Silver + Watermarking
Window Aggregations
Gold Metrics
COMPLETED
LAKEHOUSE · FRAUD & RISK ANALYTICS

Financial Fraud & Risk Lakehouse

PySpark Lakehouse applying business risk rules to synthetic financial transactions for fraud detection and customer-level risk scoring, with incremental processing and full audit/reconciliation.

Business-rule-based fraud detection and customer risk scoring across an 11-notebook Bronze / Silver / Gold Delta Lake pipeline.
PySparkDelta LakeFraud Risk RulesRisk ScoringDatabricksGitHub Actions
Transaction Source
Bronze
Silver
Fraud Detection + Risk Scoring
Gold Fraud/Risk Metrics
IMPLEMENTED
LAKEHOUSE · RETAIL ANALYTICS

Enterprise Retail Lakehouse

Databricks-oriented retail Lakehouse pipeline covering Bronze, Silver and Gold processing, incremental Delta merges, and a reusable data-quality framework for sales, customer and product analytics.

An 8-notebook, parameterized PySpark pipeline with incremental Delta merges and a configurable data-quality framework.
PySparkDelta LakeUnity CatalogDatabricksSpark SQLGitHub Actions
Bronze
Silver
Gold Aggregations
Incremental Load
Parameterized Pipeline
Get In Touch

Let's Connect

Open to Data Engineer and Azure Data Engineer opportunities.

Download CV ↓