Skip to content
View Ashok98765vvs's full-sized avatar
😇
Focusing
😇
Focusing

Highlights

  • Pro

Block or report Ashok98765vvs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Ashok98765vvs/README.md

Hi, I'm Ashok Shankarappa — Data Engineer (Fintech & Real-Time Pipelines)

LinkedIn GitHub Email Location


Who I Am

I am a Data Engineer specializing in real-time streaming pipelines, Azure-based Medallion Lakehouse architecture, PySpark distributed processing, and production-grade data quality systems.

I build the kind of data infrastructure that fintech and financial-services teams rely on to make high-stakes decisions — from sub-minute fraud alerts to $50B+ mortgage-servicing dashboards.

  • MS Computer Science — Auburn University at Montgomery (Dec 2026)
  • US Work Authorization — No sponsorship required, available immediately
  • Target roles: Data Engineer | Analytics Engineer | Data Platform Engineer

Proven Business Impact

Metric Result
Cloud compute cost reduction 35% via dbt model tuning (Snowflake/BigQuery)
Query performance improvement 45% faster average runtimes
Data incident reduction 60% fewer incidents using Monte Carlo + Great Expectations
Platform scale managed $50B+ mortgage-servicing platform (2TB → 15TB+)
Manual reporting time eliminated 6 hours → 5 minutes via Airflow automation

Featured Projects

Real-Time Azure Data Quality & Anomaly Detection Pipeline

Stack: PySpark · Azure Event Hubs · ADLS Gen2 · Delta Lake · Azure Synapse · ADF · Power BI · GitHub Actions

Production-grade pipeline ingesting real-time financial events, enforcing schema and business rules across Bronze/Silver/Gold Medallion layers, and surfacing anomalies (Z-score + IQR) on a live Power BI dashboard. Includes quarantine logic, SLA monitoring, and full CI/CD.

View Repo →


Real-Time Financial Fraud Detection Pipeline

Stack: Apache Kafka · PySpark Structured Streaming · Snowflake Dynamic Tables · Power BI · Docker

End-to-end streaming fraud detection system consuming 10K+ transactions/sec from Kafka, applying velocity checks and ML-based scoring in real time, and delivering live fraud alerts on a Power BI dashboard with sub-second latency.

View Repo →


Fintech Real-Time Loan Data Pipeline

Stack: Python · Snowflake · dbt · Apache Airflow · Great Expectations · Monte Carlo · Terraform · GitHub Actions

Production-grade mortgage loan data pipeline managing $50B+ in loan portfolios. Ingests, transforms, validates, and observes loan data across Snowflake with full SOC 2 compliance, automated data quality checks, and IaC-based deployment.

View Repo →


Retail ELT Pipeline — dbt + Snowflake

Stack: dbt · Snowflake · Python · SQL · GitHub Actions

Modular ELT pipeline for retail analytics: staging → mart layer dbt models, full test suite, and automated CI/CD deployment on every push. Demonstrates dbt best practices (sources, tests, docs, incremental models).

View Repo →


Tech Stack

Cloud & Warehouses Azure (ADLS Gen2, Synapse, ADF, Event Hubs) · Snowflake · BigQuery · Redshift · Databricks

Stream & Batch Processing Apache Kafka · PySpark · Spark Structured Streaming · Delta Lake · Medallion Architecture

Orchestration & Transformation Apache Airflow · dbt (Core & Cloud) · Azure Data Factory · Fivetran

Data Quality & Observability Great Expectations · Monte Carlo · Custom SLA Monitoring · Anomaly Detection (Z-score/IQR)

DevOps & IaC Terraform · GitHub Actions · Docker · CI/CD · Bicep

Languages Python 3.10+ · PySpark · SQL · Bash


Experience Snapshot

Bessent Technologies — Data Engineer (May 2019 – May 2020)

  • Built real-time streaming pipelines and Azure-based data platforms for financial-services clients.
  • Delivered PySpark-based ETL workflows and data quality frameworks at scale.

Let's Connect

I am actively seeking Data Engineer / Analytics Engineer / Data Platform Engineer roles in the USA. Available immediately. No sponsorship required.

Pinned Loading

  1. azure-realtime-data-quality-pipeline azure-realtime-data-quality-pipeline Public

    Production-grade Azure Data Quality & Anomaly Detection Pipeline using PySpark, Delta Lake (Medallion), Azure Synapse, ADF, and Power BI

    Python

  2. fintech-realtime-loan-pipeline fintech-realtime-loan-pipeline Public

    Real-time mortgage loan data pipeline using Python, Snowflake, dbt, Airflow, Great Expectations & Monte Carlo observability — Production-grade fintech data engineering

    Python

  3. real-time-fraud-detection real-time-fraud-detection Public

    Production-grade Real-Time Financial Fraud Detection Pipeline: Kafka, PySpark Structured Streaming, Snowflake, Power BI

    Python

  4. retail-elt-dbt-snowflake retail-elt-dbt-snowflake Public

    End-to-end Retail ELT Pipeline using dbt, Snowflake, and GitHub Actions CI/CD. Features data modeling, testing, and analytics-ready tables.