Software engineer focused on AI systems, compilers, GPU computing, and full-stack product development. I recently graduated from Tecnológico de Monterrey and enjoy building systems that connect ambitious technical ideas with reliable, measurable results.
- AI and ML systems: LLM inference, constrained decoding, structured generation, evaluation, and compiler-guided feedback loops.
- Compilers and GPU computing: MLIR/LLVM, Triton, PTX, PyTorch, XGrammar, vLLM, and CuPy.
- Backend and product engineering: Python, FastAPI, Node.js, React, Next.js, TypeScript, SQL, and REST APIs.
- Infrastructure: Docker, Kubernetes, Linux, CI/CD, Git, and Oracle Cloud Infrastructure.
An LLM-to-GPU compiler that converts grammar-constrained JSON ASTs into validated Triton MLIR, lowers kernels to PTX, and executes them against PyTorch references through a CuPy FFI runtime.
- Improved zero-shot end-to-end correctness by 17.5×, from 2.4% (4/166) to 42.2% (70/166), over direct Triton generation using the same 9B model on TritonBench-T.
- Raised PTX compilation success from 30.7% to 53.0% through a compiler-in-the-loop repair system that recovered 37 additional kernels within three attempts.
- Accepted for oral presentation at the 25th Mexican International Conference on Artificial Intelligence (MICAI 2026); proceedings forthcoming.
Built a Python data pipeline that collected and normalized metadata from more than 20 public sources for a platform used to explore and compare 300+ generative AI models. The platform included natural-language discovery powered by the OpenAI API.
Developed an IoT bicycle-management system combining GPS/RFID tracking, Firebase, Google Maps, Arduino, and an Android application.
Outside of software, I enjoy strength training, cooking, music, philosophy, and solving algorithmic problems.
