landscape · ~1200×800px
LLM Performance Benchmarking Study
A controlled evaluation framework comparing ChatGPT, Claude, DeepSeek, and Gemini across standardized programming tasks. I helped design the scoring criteria and measured output quality, accuracy, and efficiency across CS coursework and real-world software tasks, then analyzed where each model held up and where it didn't. Presented at a formal academic poster session.



