Building AI systems end to end: perception models, forecasting pipelines, and the real-time applications that have to run them.
I am an AI undergraduate at Sejong University, Department of Artificial Intelligence, graduating early in February 2027. Most of my work sits between a model and the thing that uses it β I train the model, then build the application that has to run it in real time, usually on ordinary hardware. Across all of it I care more about whether an evaluation protocol can be trusted than about the headline metric it produces.
νκ΅μ΄ μκ°
μΈμ’ λνκ΅ μΈκ³΅μ§λ₯νκ³Ό μ¬ν μ€μ΄λ©° 2027λ 2μ μ‘°κΈ°μ‘Έμ μμ μ λλ€. μμ λλΆλΆμ΄ λͺ¨λΈκ³Ό κ·Έ λͺ¨λΈμ μ°λ 물건 μ¬μ΄μ μμ΅λλ€. λͺ¨λΈμ νμ΅μν¨ λ€, κ·Έκ²μ μ€μκ°μΌλ‘ λλ €μΌ νλ μ ν리μΌμ΄μ κΉμ§ λ§λλλ€. νΉμ μ₯λΉκ° μλλΌ μΌλ°μ μΈ νλμ¨μ΄λ₯Ό μ μ λ‘ μ€κ³ν©λλ€. μ΄λ λΆμΌλ μ±λ₯ μμΉ μ체보λ€, κ·Έ μμΉλ₯Ό λ§λ€μ΄λΈ νκ° λ°©λ²μ μ λ’°ν μ μλμ§λ₯Ό λ¨Όμ νμΈν©λλ€.
| λΆμΌ | λ΄μ© |
|---|---|
| λ©ν°λͺ¨λ¬ μΈμ§ | μμ Β·μμΈΒ·νκ²½μμ ν¨κ» λ°μ κ³ μ μκ³κ°μ΄ μλ μ¬μ©μλ³ κΈ°μ€μ μΌλ‘ νμ |
| μκ³μ΄ μμΈ‘ | 물리μ κ·Όκ±°λ₯Ό κ°μ§ νΌμ²λ‘ λ€μ€ horizon νκ·. μμ νΌμ² μ§ν©μ΄ μ΄λ―Έμ§ μΈμ½λλ₯Ό μ΄κΈ°λ κ²½μ°λ₯Ό λ€λ£Έ |
| μ€μκ° μμ€ν | WebSocketΒ·WebRTC νμ΄νλΌμΈ β μ격 κ°λ , λ€μ€ μ°Έκ°μ λκΈ°ν, νμ νλ©΄Β·μ€λμ€ μ·¨λ |
| LLM μμ© | Bedrock κ²½μ Claudeλ‘ λ΄λ μ΄μ μμ± λ° λ³ν λ¬Έμ μΆμ νμ΄νλΌμΈ κ΅¬μ± |
| λ°μ€ν¬ν±Β·μλν | Electron μ ν리μΌμ΄μ , κ·Έλ¦¬κ³ APIκ° μλ λ¨λ§μμ GUI κ³μΈ΅λ§μΌλ‘ μννλ μ 무 μλν |
SEE-ON β λ°νμκ° λ§νμ§ μμ μκ° μ 보λ₯Ό νλ©΄ ν΄μ€ λ΄λ μ΄μ μΌλ‘ λ§λ€μ΄ νμ νλ¦μ λμ§ μλ μ§μ μ λ£λ λ©ν° μμ΄μ νΈ μμ€ν μ λλ€. AI Rookie λ³Έμ μ§μΆμμ΄λ©° νμ₯μ λ§‘κ³ μμ΅λλ€. See-on26 μ‘°μ§μ λ κ°μ λΉλκ° κ³΅κ°λμ΄ μμ΅λλ€ β seeon-meetμ μ€μ Google Meet νμμμ νλ©΄κ³Ό μ€λμ€λ₯Ό μ·¨λνκ³ , electron-seeon-experimentλ λ Ήν μμμ λμΌν λ΄λ μ΄μ 체μΈμ ν¬μ νλ μ€νμ€ λΉλμ λλ€.
| Area | Focus |
|---|---|
| Multimodal perception | Gaze, pose, and ambient audio fused into a per-user baseline rather than a fixed threshold |
| Time-series forecasting | Multi-horizon regression on physically grounded features, including the cases where a small feature set beats an image encoder |
| Real-time systems | WebSocket and WebRTC pipelines β remote invigilation, synchronised multiplayer, live meeting capture |
| LLM applications | Narration and variant-question pipelines running on Claude via Amazon Bedrock |
| Desktop & automation | Electron applications, and GUI-layer automation for terminals that expose no API |
| Area | Tools |
|---|---|
| Languages | |
| ML / DL | |
| Vision / Audio | |
| LLM / Speech | |
| Backend | |
| Frontend | |
| Desktop / Infra |
SEE-ON β a multi-agent system that narrates the visual information a speaker leaves unsaid during a live video meeting. Team lead, AI Rookie finals. Two builds are public under the See-on26 organisation: seeon-meet captures screen and audio from a live Google Meet session, and electron-seeon-experiment replays recordings through the same narration chain for evaluation.