I am a Data Engineer specializing in real-time streaming pipelines, Azure-based Medallion Lakehouse architecture, PySpark distributed processing, and production-grade data quality systems.
I build the kind of data infrastructure that fintech and financial-services teams rely on to make high-stakes decisions — from sub-minute fraud alerts to $50B+ mortgage-servicing dashboards.
- MS Computer Science — Auburn University at Montgomery (Dec 2026)
- US Work Authorization — No sponsorship required, available immediately
- Target roles: Data Engineer | Analytics Engineer | Data Platform Engineer
| Metric | Result |
|---|---|
| Cloud compute cost reduction | 35% via dbt model tuning (Snowflake/BigQuery) |
| Query performance improvement | 45% faster average runtimes |
| Data incident reduction | 60% fewer incidents using Monte Carlo + Great Expectations |
| Platform scale managed | $50B+ mortgage-servicing platform (2TB → 15TB+) |
| Manual reporting time eliminated | 6 hours → 5 minutes via Airflow automation |
Stack: PySpark · Azure Event Hubs · ADLS Gen2 · Delta Lake · Azure Synapse · ADF · Power BI · GitHub Actions
Production-grade pipeline ingesting real-time financial events, enforcing schema and business rules across Bronze/Silver/Gold Medallion layers, and surfacing anomalies (Z-score + IQR) on a live Power BI dashboard. Includes quarantine logic, SLA monitoring, and full CI/CD.
Stack: Apache Kafka · PySpark Structured Streaming · Snowflake Dynamic Tables · Power BI · Docker
End-to-end streaming fraud detection system consuming 10K+ transactions/sec from Kafka, applying velocity checks and ML-based scoring in real time, and delivering live fraud alerts on a Power BI dashboard with sub-second latency.
Stack: Python · Snowflake · dbt · Apache Airflow · Great Expectations · Monte Carlo · Terraform · GitHub Actions
Production-grade mortgage loan data pipeline managing $50B+ in loan portfolios. Ingests, transforms, validates, and observes loan data across Snowflake with full SOC 2 compliance, automated data quality checks, and IaC-based deployment.
Stack: dbt · Snowflake · Python · SQL · GitHub Actions
Modular ELT pipeline for retail analytics: staging → mart layer dbt models, full test suite, and automated CI/CD deployment on every push. Demonstrates dbt best practices (sources, tests, docs, incremental models).
Cloud & Warehouses Azure (ADLS Gen2, Synapse, ADF, Event Hubs) · Snowflake · BigQuery · Redshift · Databricks
Stream & Batch Processing Apache Kafka · PySpark · Spark Structured Streaming · Delta Lake · Medallion Architecture
Orchestration & Transformation Apache Airflow · dbt (Core & Cloud) · Azure Data Factory · Fivetran
Data Quality & Observability Great Expectations · Monte Carlo · Custom SLA Monitoring · Anomaly Detection (Z-score/IQR)
DevOps & IaC Terraform · GitHub Actions · Docker · CI/CD · Bicep
Languages Python 3.10+ · PySpark · SQL · Bash
Bessent Technologies — Data Engineer (May 2019 – May 2020)
- Built real-time streaming pipelines and Azure-based data platforms for financial-services clients.
- Delivered PySpark-based ETL workflows and data quality frameworks at scale.
I am actively seeking Data Engineer / Analytics Engineer / Data Platform Engineer roles in the USA. Available immediately. No sponsorship required.
- Email: ashokchowdary776@gmail.com
- LinkedIn: linkedin.com/in/ashok-s1
- GitHub: github.com/Ashok98765vvs
