Software Expert Engineer · MercorContract
Apr 2025 — PresentRemote (Worldwide)
- Validated and adapted 50+ open-source GitHub repositories to evaluate model behavior for frontier LLMs (Grok 4, Claude); resolved environment configuration issues across Python, SQL, MySQL, PostgreSQL, Docker, and Podman.
- Designed prompt evaluation rubrics and documented model failure modes, contributing to measurable robustness improvements in AI model performance.
- Reviewed and evaluated pull requests for correctness, test coverage, and code quality across a distributed team workflow.