# Evaluation: Aurora Labs — Senior AI Platform Engineer

**Date:** 2026-07-11
**Archetype:** AI Platform / LLMOps Engineer
**Score:** 4.2/5
**URL:** https://demo.jobs.clawvard.school/aurora-labs/senior-ai-platform-engineer (synthetic demo posting)
**Legitimacy:** trusted (demo)
**PDF:** output/cv-alex-chen-aurora-labs-2026-07-11.pdf
**Cover letter:** output/aurora-labs-cover-letter.pdf

> Demo posting used for the Clawvard `career-ops-copilot` course. Aurora Labs is a synthetic employer and Alex Chen is a synthetic candidate — no real hiring pipeline, comp data, or personally-identifying information is referenced.

---

## 1) 角色摘要

| 字段 | 值 |
|-------|-------|
| **Archetype** | AI Platform / LLMOps Engineer |
| **领域** | Platform / Infrastructure |
| **职能** | Build |
| **级别** | Senior (IC4–IC5) |
| **远程** | Full remote (US timezone overlap required) |
| **团队规模** | ~7 platform engineers（来自 JD） |
| **一句话** | Senior AI eng to build and scale LLM serving + eval infrastructure for Aurora's enterprise customers, own the rollout gate and observability stack. |

## 2) 简历匹配

| JD Requirement | CV Match | Source |
|----------------|----------|--------|
| "Production LLM systems, not just fine-tuning notebooks" | Built LLM reranking on top of collaborative filtering, 18% conversion uplift over 12 A/B weeks | cv.md: TechFin Corp |
| "Model rollout with confidence — eval gates, not offline reports" | Wired CI/CD with automated eval gates against held-out set; deploy time 2 weeks → 4 hours | cv.md: TechFin Corp |
| "Observability: drift, latency, and cost as first-class SLIs" | Stood up drift dashboards + performance SLIs + retraining triggers; on-call load −35% | cv.md: TechFin Corp |
| "Python + distributed systems (Kafka / K8s / feature store)" | Kafka streaming pipeline at 50 ms p99, feature store, Kubernetes | cv.md: TechFin Corp + Skills |
| "Open-source ML tooling" | Maintains LLM Eval Toolkit + FraudShield reference stack | cv.md: Projects |

### 差距

| Gap | Severity | Mitigation |
|-----|----------|------------|
| "Multi-tenant LLM serving at 10⁵ RPS" | Medium | Frame TechFin serving stack as the same shape at fintech scale; call out the per-tenant rate limiter design in the interview loop. |
| "Cost-per-successful-request SLIs" | Low | Explicit call-out in cover letter — makes the gap read as intentional prioritization, not blind spot. |

## 3) 级别与策略

**Detected level:** Senior (IC4)
**Candidate's natural level:** Senior–Staff boundary.

**"Sell senior" plan:** Lead the recruiter screen with "ML platform team lead, 3 direct reports, 4 downstream product teams." Frame the eval-gate + monitoring work as the platform-ownership signal, not the reactive on-call story. Ready for Staff scope conversations if they open.

## 4) 薪资与需求

| Data Point | Value | Source |
|------------|-------|--------|
| Base salary range (public demo) | $185–225K | Public comp-band sources referenced in JD demo |
| Total comp (with equity) | $260–330K | Public range for equivalent Senior AI Platform role in the demo |
| Demand trend | High — LLM infra remains top-5 in-demand IC role | Public trend data |

> Comp figures above are illustrative for the course demo — the real career-ops `oferta` mode fetches them per posting through the user's own configured research budget.

## 5) 定制方案

| # | 段落 | 现状 | 建议改动 | 原因 |
|---|---------|---------|-----------------|-----|
| 1 | Summary | "Full-stack ML engineer" | "Senior AI engineer focused on LLM infrastructure, evaluation, and observability" | Mirror JD language exactly. |
| 2 | TechFin bullets | Generic ML platform framing | Lead with LLM-serving stack + rollout gates | Aurora's JD explicitly names rollout confidence as the pain. |
| 3 | Projects | Both listed equally | Lead with LLM Eval Toolkit | Direct evidence of eval-first mindset. |
| 4 | Skills | Grouped by generic tag | Move "LLM eval / prompt engineering" up alongside PyTorch | ATS keyword mirroring. |

## 6) 面试准备（情景 / 任务 / 动作 / 结果 / 反思）

| # | JD Requirement | 故事 | 情景 | 任务 | 动作 | 结果与反思 |
|---|---------------|------------|---|---|---|----|
| 1 | Production LLM systems | LLM reranking on collaborative filtering | Legacy recsys plateaued at 12% CTR; recs felt "same-y" | Add LLM reranker without regressing latency SLO | Built lightweight reranker, tuned per-tenant, gated via A/B with automated rollback | 18% conversion uplift over 12 A/B weeks. 反思：the win came from the eval gate, not the model — I'd design that first next time. |
| 2 | Rollout with confidence | Eval-gate CI/CD | Deploys took 2 weeks and shipped 1-in-4 regressions | Cut cycle time without losing safety | Wired GitHub Actions + SageMaker + held-out eval gate that blocks merges | 2 weeks → 4 hours; regressions caught before prod. 反思：the missing rung was per-slice eval — global metrics hid a segment regression once. |
| 3 | Observability & on-call | Drift + perf dashboards | Model monitoring was tail-drop metrics only | Get ahead of drift before customers noticed | Stood up drift dashboards, perf SLIs, automated retrain triggers on breach | On-call load −35%; time-to-detect drift 5 days → 6 hours. 反思：I'd add cost/req as a peer SLI on day one — that's still my open ticket. |

**Recommended case study:** LLM Eval Toolkit — direct evidence of the eval-first mindset the JD calls out, plus open-source impact.

---

## Keywords Extracted

LLM infrastructure, model serving, eval gates, rollout confidence, observability, drift monitoring, cost per request, Python, Kubernetes, Kafka, feature store, multi-tenant serving, per-tenant rate limiting, production LLM, distributed systems.

---

## 岗位真实性检查

| 信号 | 结论 | 备注 |
|--------|---------|------|
| 公开雇主 | ✅ 合成 demo — 无真实招聘风险 | 课程展示专用 |
| 薪资透明度 | ✅ 明确带宽 | 与公开数据一致 |
| 角色描述特异度 | ✅ 明确指向 LLM 基础设施，不是 "AI 通才" | 非幽灵 / 喷投挂牌 |
| 联系人清晰度 | ✅ 明确团队 + 投递路径 | (合成 demo) |
| 90 天重复挂牌信号 | ✅ 首次出现 | (合成 demo) |

**真实性等级：** `trusted` (demo)。真实运行时同一检查会标出幽灵岗位、中介转发或 90 天内重复挂牌 2 次以上的岗位。
