Incoming second-year undergraduate at SJTU SPEIT | LLM & software-agent evaluation · interested in embodied intelligence & robotics
Undergraduate entering my second year at Shanghai Jiao Tong University, enrolled in the Paris Elite Institute of Technology (SPEIT). My prior research experience is in LLM evaluation: building evaluation benchmarks for LLM agents and studying structured-output robustness. I am now actively expanding toward embodied intelligence and robotics, while continuing to explore AI evaluation.
- Large language models(大语言模型)
- Embodied intelligence & robotics(具身智能与机器人学)
- Mechanical systems(机械)
- Did It Happen? Counterfactual Evaluation of LLM Agent Recovery from Ambiguous Tool Outcomes — Research Square preprint, 2026. DOI
- Mutation Testing of Task-Scoped State Oracles in Software-Agent Benchmarks: A Cross-Benchmark Empirical Study — Research Square preprint, 2026. DOI
- Testing JSON Schema Instruction Artifacts: Distributional Robustness under Validation-Equivalent Serialization and JSON Mode — Research Square preprint, 2026. DOI
- Python — basic working proficiency(具备基础实用能力)
- AI-native workflows — comfortable driving AI-assisted research and engineering workflows(AI 原生工作流)
Research:
- Ambiguous Tool Outcomes Benchmark — counterfactual benchmark for LLM agent recovery from ambiguous tool outcomes
- Side-Effect Calibration Study — mutation testing of task-scoped state oracles in software-agent benchmarks
- Schema Order Robustness — JSON Schema serialization-order robustness in black-box LLM generation
Other:
- VEX Robotics — VEX V5 robot control code; lead programmer, 2nd Prize at SJTU campus competition
- Snowbound School Mystery(雪闭校园) — Ren'Py suspense visual novel, 3 branching routes · play online
B.Eng. — Shanghai Jiao Tong University, Paris Elite Institute of Technology (SPEIT) 2025 – Present · Shanghai, China
- Email: sthfornothing@sjtu.edu.cn
- ORCID: 0009-0008-9175-8226
- GitHub: @shushuyang231
- Academic Page: shushuyang231.github.io