Skip to content
View YusefSyed's full-sized avatar

Highlights

  • Pro

Block or report YusefSyed

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
YusefSyed/README.md

Yusef Syed. Founder of Scopehaven. Testing what AI agents can access and do. Student at the University of Toronto.

Scopehaven Portfolio LinkedIn Email me Resume PDF

Hey, I'm Yusef. My main focus is building Scopehaven: permission and outcome testing for AI apps and agents. I'm also a first-year student at the University of Toronto, intending Math + CS.

What I'm working on

  • 🔍 Scopehaven: checking who can access data or use a tool, comparing the agent's reply with observable outcomes, and rerunning checks after a repair. Current evidence comes from local tests and internal pilots with synthetic data. The product is early stage, and the code is private. See the approach and request a first test.
  • 🧪 Agent evaluation: my eval lab studies tool failures, partial execution, and missing evidence. I'm interested in evaluations that distinguish an intended action from a verified result.
  • 🧩 Open source: contributing focused fixes to evaluation and research tools, with regression tests and links to the upstream review.

What I'm working toward

  • Make Scopehaven useful to external teams: reproduce a concrete permission failure, verify a repair, and preserve the check for future releases.
  • Build evaluations that catch forbidden actions while checking that legitimate use still works, and report uncertainty when the evidence is incomplete.
  • Strengthen my independent Python, algorithms, debugging, and mathematical foundations alongside my coursework.

💻 Selected projects & experiments

Scopehaven: tests agent permissions and outcomes, then reruns checks after a fix. Early-stage product with local and synthetic pilot evidence.

agent-eval-mutation-lab: tests agent actions, partial failures, and missing results. Python. Providence: my native Apple Watch client for our Hack the North team project. Swift. tiraz-garment-completion: annotation-only PyTorch experiment with calibration and missing-context tests. agent-proof: reviewer-selected checks and redacted reports. TypeScript CLI. shiftproof: local scheduling prototype with independent constraint checks. Python. callreclaim-webmcp: a synthetic missed-call demo where the owner decides. TypeScript.

See all my repositories

🧩 Work I've contributed to

Microsoft Agent Lightning: merged shutdown race fix. Inspect Scout: merged model-usage accounting fix. Kornia: merged empty accelerator tensor fix. Microsoft PyRIT: merged package-hallucination techniques. WandB RAI Toolkit: merged default scorer category fix. PyTorch: open CPU/CUDA incomplete-gamma shape-gradient PR, not yet merged.

Each card links to my contribution. Green labels indicate merged PRs; the PyTorch PR is still open as of October 8, 2026.

My merged work includes Apache Arrow slicing, Agent Lightning timeout-report retries, Inspect Scout postponed-annotation support, and W&B RAI Toolkit prediction-error accounting. The evidence page lists the verified contributions and the date checked.

Contribution details and evidence

📱 My apps

The public CallReclaim Agent Desk is a separate synthetic demo with no messaging backend. ShiftProof is also a local prototype using synthetic data.

🛠️ What I work with

Languages: Python · TypeScript · JavaScript · Swift · SQL

Apps: React · React Native · Expo · Next.js · SwiftUI · watchOS

Experiments: PyTorch · NumPy · Inspect · pytest

Infrastructure: PostgreSQL · SQLite · Supabase · Docker · GitHub Actions

🔬 Research notes

The eval lab includes a frozen 624-trial local-model study. Invalid-output rates differed between conditions, so I withheld improvement claims and kept the descriptive results and missing-output analysis.

The Tiraz study uses garment annotations, not images. Its three seeded models achieved 58.5–59.0% top-1 against a 52.7% baseline on 962 held-out groups; removing context exposed a large drop in prediction-set coverage.

Methods, results, and limitations · PyTorch local validation supplement

📬 Get in touch

If you're building an AI app or agent and want to discuss permission testing, evaluation failures, or a focused open-source contribution, you can reach me below.

Résumé · Email · LinkedIn

Typing intro and compact grids inspired by DenverCoder1.

Pinned Loading

  1. agent-eval-mutation-lab agent-eval-mutation-lab Public

    Reproducible experiments that check what tool-using agents actually did, including partial failures and unknown outcomes.

    Python

  2. tiraz-garment-completion tiraz-garment-completion Public

    A PyTorch garment-completion study measuring accuracy, calibration, and uncertainty when outfit context is missing.

    Python

  3. agent-proof agent-proof Public

    Run reviewer-selected checks and keep redacted, traceable evidence. A dependency-free TypeScript CLI.

    TypeScript

  4. callreclaim-webmcp callreclaim-webmcp Public

    A WebMCP missed-call desk where agents prepare replies and owners decide. Interactive demo with synthetic data.

    TypeScript

  5. shiftproof shiftproof Public

    A local volunteer-scheduling agent with independent constraint checks and coordinator approval. Synthetic-data prototype.

    Python