San Francisco, CA — open to relocation
Joseandres Hinojoza
Software Engineer / Site Reliability Engineer
Site Reliability Engineer and Software Engineer with 4+ years of experience building and operating distributed systems at scale. Spent 2.5 years at Google on a team operating services at 150M+ QPS, designing a monitoring system that covered 300+ clusters and analyzed 10M+ QPS of health-check traffic. Has since built observability infrastructure for LLM inference and contributed to frontier model evaluation pipelines. Focused on distributed systems, incident response, and production reliability.
2.5 yrs
Site Reliability Engineer at Google
150M+
QPS handled by team-owned services at Google
300+
Clusters covered by monitoring systems built
4+ yrs
Professional software engineering experience
Recent experience
Software Expert Engineer · Mercor
Apr 2025 — PresentRemote (Worldwide)
- Validated and adapted 50+ open-source GitHub repositories to evaluate model behavior for frontier LLMs (Grok 4, Claude); resolved environment configuration issues across Python, SQL, MySQL, PostgreSQL, Docker, and Podman.
- Designed prompt evaluation rubrics and documented model failure modes, contributing to measurable robustness improvements in AI model performance.
Software Engineer Intern · Near AI
Nov 2025 — Dec 2025San Francisco, CA, United States
- Designed and deployed a monitoring system for 5 open-source LLMs served via vLLM; implemented health probes, alerting pipelines, and real-time dashboards to ensure high availability and model reliability.
- Configured and optimized NGINX-based load balancing across multiple inference services, improving traffic distribution and reducing tail latency.
Core skills
Languages
- Python
- C++
- C
- Rust
- Java
- JavaScript
- Bash
- SQL
Infrastructure & Cloud
- Kubernetes
- Docker
- Podman
- Terraform
- Ansible
- Borg
- GCP
- AWS
Observability & Reliability
- Prometheus
- Grafana
- Datadog
- PagerDuty
- Monarch
- SLOs / SLIs / Error Budgets
- Chaos Engineering
- Canary & Blue/Green Deployments