MS in Data Science, Rochester Institute of Technology (May 2026). I design and build data pipelines end to end — streaming and batch architectures, ETL, and applied ML — and back them with tests and CI rather than just a demo script.
Reach me: soorajmanoj@gmail.com · LinkedIn · Portfolio
Graduate Teaching Assistant, Department of Computer Science — Rochester Institute of Technology (Aug 2025 – Apr 2026) Mentored 50+ students in server programming and backend system design; reviewed student-built data pipelines for correctness and efficiency.
Intern, Analysis & Design of Commissions for Gig Workers — One Integral Technologies Pvt. Ltd. (2022 – 2024) Built performance-based pay models and a React.js analytics dashboard on top of a MongoDB-backed data pipeline serving 500+ gig workers.
pyspark-etl-healthcare-normalization-pipeline
PySpark ETL pipeline that normalizes flat legacy healthcare data into a real 10-dimension + 1-fact snowflake schema, with automated referential integrity checks across every foreign key.
credit-card-kafka-pipeline
Real-time + batch credit card transaction processing system on a Lambda architecture — Kafka stream layer for provisional fraud checks, batch layer for reconciliation and credit scoring, MySQL for persistence.
youtube-ai-assistant
Ask questions about any YouTube video's content — transcript extraction, FAISS vector search, and a local LLM (Ollama) for retrieval-augmented Q&A, no external API keys required.
dockerized-mlops-api-fastapi
Containerized FastAPI service serving a scikit-learn (TF-IDF + Logistic Regression) sentiment classifier via Docker, with pytest coverage and CI that validates both the test suite and the Docker build.
Bias in AI-Generated Travel Narratives: Evaluating LLMs in Countering Misconceptions — DSCI 602 final paper, RIT Co-authored with Anshuman Mohapatra, Bhaskar Sreenivas Pavan Akkena, and advisors Ashique Khudabukhsh and Travis Desell. A 5-layer pipeline collecting ~600K YouTube travel comments, generating counterspeech across four LLMs (LLaMA32, Gemma2, GPT2, Sarvam), and comparing human toxicity ratings against Detoxify scores — finding that automated toxicity classifiers systematically underestimate human-perceived harm and produce unstable model safety rankings.
- Microsoft Azure AI Fundamentals (AI-900)
- Diploma in Programming and Data Science, IIT Madras
- Six Sigma White Belt (Council for Six Sigma Certification)

