Associate Software Engineer at Accenture with hands-on experience in Azure Databricks, Apache Spark, and building batch & streaming ELT pipelines using PySpark and Delta Lake. Oracle-certified Java developer with a B.Tech in Computer Science and Business Systems — passionate about scalable data platforms and cloud technologies.
</About Me>
Hi, I'm Balaganesh S B — Associate Software Engineer at Accenture, working at the intersection of data engineering and cloud technology. I hold a B.Tech in Computer Science and Business Systems from Panimalar Engineering College (CGPA: 8.97, 1st Class with Distinction).
I specialize in building batch and streaming ELT pipelines using Azure Databricks, Apache Spark (PySpark & Spark SQL), and Delta Lake, following Medallion Architecture principles (Bronze → Silver → Gold). I'm an Oracle-certified Java developer with solid foundations in Spring Boot, Angular, and full-stack development — gained through my internship at Wipro TalentNext.
Beyond code, I express creativity through photography, capturing stories in everyday moments. I also enjoy movies and fitness — balancing a tech-driven mind with an active lifestyle.
</Experience>
- Azure Databricks & Apache Spark: Completed role-based training on Azure Databricks, building batch and streaming ELT pipelines using PySpark, Spark SQL, and Delta Lake following Medallion Architecture (Bronze → Silver → Gold) principles.
- Data Engineering Solutions: Currently contributing to enterprise data engineering solutions on cloud platforms, designing scalable pipelines for large-scale data ingestion and transformation.
- Structured Streaming & Delta Lake: Gained hands-on experience with real-time data processing using Structured Streaming and ACID-compliant Delta Lake tables for reliable data lakes.
</Skills>
Data Engineering Stack
Azure Databricks
Apache Spark
Delta Lake
SQL
Full-Stack & Mobile
Java
Android SDK
Spring Boot
Angular
React
SQLite
HTML5
CSS3
Bootstrap
JavaScript
Python
C
Tools & DevOps
Git
GitHub
Docker
</Certifications>
</Projects>
Enterprise Retail Lakehouse Platform
Built an end-to-end retail lakehouse on Databricks using the Medallion architecture, Auto Loader, Delta Live Tables, Structured Streaming, and Databricks SQL for scalable analytics.
Metadata-Driven-ETL-Framework
Developed a metadata-driven ETL framework that automates configurable data ingestion, SCD processing, and data quality using generic pipelines.
Banking Streaming Analytics Platform
Built a real-time banking analytics platform for streaming transaction processing, fraud detection, and KPI monitoring.
SF Fire Lakehouse Pipeline
Built an end-to-end Lakehouse pipeline on Databricks using 4.38M+ San Francisco Fire Department records. Implemented batch, streaming (Auto Loader), and Delta Live Tables pipelines using the Medallion Architecture for scalable and optimized data processing.
Food Delivery Data Pipeline
Developed batch and streaming data pipelines for food delivery datasets using PySpark and Delta Lake. Implemented Auto Loader, incremental MERGE/UPSERT processing, and SQL analytics to deliver analytics-ready data.