Good evening, Sergii 👋
You have 7 active campaigns · 2 need attention · $314 spent on Acme Q2 this week
Acme Q2 Launch
Needs attention client:Acme · project:Q2B2B fintech product launch · 12 graphs · DAG with 18 steps · scheduled daily 09:00 UTC
| Time | Status | Steps | Cost | Duration |
|---|---|---|---|---|
| 14:32 | success | 18/18 | $1.20 | 12.4s |
| 14:18 | partial | 12/18 | $0.40 | 47.1s |
| 14:02 | failed | 4/18 | $0.05 | 3.2s |
| 13:30 | success | 18/18 | $1.18 | 11.1s |
| 13:01 | success | 18/18 | $1.22 | 12.8s |
| 12:30 | success | 18/18 | $1.15 | 10.9s |
tone_classifier on branch mainBrand Guides re-indexed (1.2K chunks)tone_classifier @ stricter-rubricQuality dropped 17pp on tone_classifier after a prompt edit at 14:10. Eval suite was not run.
🔍 Inspector Langfuse Grafana Cloud Traces
142 runs in the last 7 days · 3 failed · $314 cost · scope: Acme Q2 Launch
generate_promo_text click any step on the left to switch{
"headline": "Q2 brings precision to your treasury.",
"bullets": [
"Real-time cash visibility for CFO teams.",
"Built for regulated industries — SOC2, HIPAA-ready.",
"Onboard in days, not quarters."
],
"cta": "Book a 15-min demo"
}compose_email, compose_landing_html, generate_dalle_heroYou are a copywriter for Acme Bank. Tone: professional, warm. Brand context (KB): [3 chunks · 1,247 t] Brief: Q2 launch for Acme Bank's new B2B fintech product… Output a punchy headline + 3 supporting bullets in JSON.
{
"headline": "Q2 brings precision to your treasury.",
"bullets": [
"Real-time cash visibility for CFO teams.",
"Built for regulated industries — SOC2, HIPAA-ready.",
"Onboard in days, not quarters."
],
"cta": "Book a 15-min demo"
}
🔀 Workflow Editor
Acme Q2 Launch · generate_promo_assets · 18 nodes · cost preview $1.20/run · 12.4s avg
🔒 14 sealed 🔓 3 broken (in CP-104) — 1 unsealed (TRUST tier) 🕸 see blast radius for this DAG 📋 active Change Plan
- generate_hero (MULTI_LAYER) — moderate_image is connected, but no human_approval edge before publish · add edge
- extract_intent outputs
{intent, audience}but downstreamapply_toneonly consumestone— unused output · view types
generate_hero can publish without human approval). Auto-fix is available: insert human_approval node between moderate_image and publish.{{brief}}from fetch_brief{{kb_results}}from kb_lookupintent : enum→ apply_toneaudience : stringunused ⚠🧠 Model Catalog
9 models from 4 providers · 2 BYOK keys configured · 1 deprecation upcoming
| Model | Status | Cost (in/out) | p95 latency | Context | Used by | Calls 7d |
|---|---|---|---|---|---|---|
| claude-3-5-sonnet-20241022 default · review | ● active | $3 / $15 | 920ms | 200K | 14 prompts | 8.2K |
| claude-3-5-haiku-20241022 default · draft | ● active | $0.80 / $4 | 410ms | 200K | 5 prompts | 2.1K |
| claude-3-opus-20240229 final / sunset Q3 | ● deprecating Jul 31 | $15 / $75 | 2.1s | 200K | 2 prompts | 340 |
| Model | Status | Cost (in/out) | p95 latency | Context | Used by | Calls 7d |
|---|---|---|---|---|---|---|
| gpt-4o-2024-11-20 fallback | ● active | $2.50 / $10 | 1.4s | 128K | via Always-On | 421 |
| gpt-4o-mini judge | ● active | $0.15 / $0.60 | 680ms | 128K | eval judges | 214 |
| dall-e-3 | ● active | $0.04 / image | 3.2s | — | 1 prompt | 26 |
| Model | Status | Cost (in/out) | p95 latency | Context | Used by | Calls 7d |
|---|---|---|---|---|---|---|
| gemini-1.5-pro | ⚠ degraded | $1.25 / $5 | 4.2s | 2M | 1 prompt | 12 |
| gemini-1.5-flash | ● active | $0.075 / $0.30 | 540ms | 1M | — | 0 |
💬 Customer Feedback Grafana Cloud Metrics
312 signals this week · NPS 47 (▲ 4) · 14 new in queue · 3 converted to eval cases
tone_classifier rev 142. Suggested action: view Quality Insights alert · or open Hypothesis "tone calibration for B2B finance".| Campaign | 👍 | 👎 | Ratio | Trend |
|---|---|---|---|---|
| Delta Newsletter | 42 | 2 | 95% | ▇▇▇▇▇▇ |
| Bravo Spring | 38 | 5 | 88% | ▆▇▆▇▇▇ |
| Acme Q2 | 52 | 12 | 81% ▼ | ▇▇▇▆▅▄ |
| Charlie Holiday | 28 | 3 | 90% | ▇▇▇▇▇▇ |
| Echo Black Friday | 21 | 8 | 72% ▼ | ▆▅▄▄▃▃ |
Drop this snippet on customer-facing surface (email, landing, chat) to collect feedback:
<script src="https://cdn.velocityengine.io/feedback.js" data-workspace="acme-marketing" data-campaign="acme-q2"></script>
📈 Quality Insights Grafana Cloud · live
Acceptance 87.3% · 2 open alerts · 47 rejected runs since 14:48
| Time | Run | Claim | Source check | Action |
|---|---|---|---|---|
| Apr 18 16:22 | #4712 | "Acme Bank holds $2.3B AUM" | not in GC · not in brief | blocked |
| Apr 17 11:09 | #4604 | "audited by Big Four" | contradicts brief: "Acme is private" | blocked |
| Apr 16 09:44 | #4521 | "40% YoY growth" | brief says "20% YoY" | blocked |
| Apr 15 14:31 | #4480 | "NYSE-listed" | not stated · public web check fails | blocked |
| KB / Source | Last sync | Source updated | Affected campaigns | Action |
|---|---|---|---|---|
brand_guides rev 7 | Mar 4 | Apr 12 · 6 sections changed | Acme Q2, Bravo Spring | Re-sync |
compliance_rules_eu rev 12 | Feb 28 | Apr 10 · GDPR addendum | Charlie Holiday EU | Re-sync |
tone_examples rev 4 | Apr 1 | Apr 18 · Jess added 12 | Echo Black Friday | Re-sync |
product_catalog rev 22 | Apr 17 | Apr 19 · 3 SKUs added | Foxtrot Loyalty | Re-sync |
chat_assistant_acme · 87% acceptchat_holiday_promo · 64% (needs attention)compose_subject_v2 · 28% edit-then-applyPrompt tone_classifier was edited 38m before the drop.
Inputs to this prompt drifted: queries about "compliance" rose 3.2× this week vs eval set.
→ Add cases to eval suite→ See related Customer Feedback (3 patterns)
⚡ Always-On AI Portkey Grafana Cloud SLO · OnCall · Incident
4 policies active · Anthropic + OpenAI healthy · Google AI degraded · 47 fallbacks today (+$8.40)
Multi-provider HA policy as Tier 2. Either drop it from chain until recovery, or add Bedrock as backup tier so HA stays redundant. Live calls already auto-routing to other tiers.| Time | Event | Provider | Detail | Run |
|---|---|---|---|---|
| 14:42 | ⇄ fallback | Anthropic → OpenAI | 5xx (Sonnet 503) · 3 retries exhausted | #4828 |
| 13:18 | ↻ retry | Anthropic | rate_limit 429 · backoff 500ms | #4819 |
| 12:42 | ⇄ fallback | Anthropic → OpenAI | 5xx burst · 3 retries exhausted (today's primary incident) | #4811 |
| 11:00 | ⊘ CB open | Google AI | failure rate 67% > 50% threshold · open 60s | — |
| 10:15 | ↻ retry | OpenAI | schema_violation · re-prompt | #4799 |
Dry-run a failure to see what your policy would do.
✏️ Prompt Studio PromptLayer
Editing promo_copywriter @ experiment-tone-v2 · 2 linter warnings · cost preview $0.018/call on Sonnet
promo_copywriter, compose_email
🕸 see 11 downstream
experiment-tone-v2, 2 linter warnings open, and Quality Gate hasn't run yet on this revision. Apply linter fixes + run eval against promo_copy_quality (47 cases · ~$0.85) before requesting review.{
"headline": "Yo, Acme Bank's got your treasury covered.",
"bullets": [
"See cash in real-time, anytime.",
"Compliance-ready out the box.",
"Up and running in days."
]
}
{
"headline": "Q2 brings precision to your treasury.",
"bullets": [
"Real-time cash visibility for CFO teams.",
"Built for regulated industries — SOC2, HIPAA-ready.",
"Onboard in days, not quarters."
]
}
compose_email calls headline.toLowerCase() (expects string, not array). Will throw TypeError at runtime in <eval> node.promo_copywriter:brand_tone_querybrief_clarificationcopy_revisionTypeError: x.toLowerCase is not a function in <eval> nodes downstream.{
"headline": "string · ≤80",
"bullets": "array<string> · ≥3",
"cta": "string · nullable"
}
compose_email · expects headline.toLowerCase()compose_landing_html · iterates bullets[]moderate_image · null-checks cta🚦 Quality Gate Braintrust
5 eval suites · 92% avg pass rate · 1 suite blocking publish (tone_classifier · 87% ▼)
b2b-fintech-tone
| Type | Config | Weight | Last result | |
|---|---|---|---|---|
| OUTPUT_CONTRACT | JSON · promo.schema.json · runtime: block+retry | 5 | ✗ 2/47 type-mismatch | ⋯ |
| CONTAINS | "Acme" | 1 | ✓ pass | ⋯ |
| MAX_LENGTH | 280 chars | 1 | ✓ pass | ⋯ |
| JSON_SCHEMA | promo.schema.json | 2 | ✓ pass | ⋯ |
| LLM_JUDGE | "tone matches professional+warm rubric" | 3 | ✗ 3/10 | ⋯ |
headline.toLowerCase()-style downstream consumption. Failures here block publish regardless of pass-rate (would have caught MCT key-error TypeError before reaching prod).
Convert any production run into an eval case from Inspector or directly from Customer Feedback (👎 with comment → REGRESSION case).
Record real production LLM calls into immutable cassettes; replay against any prompt branch / model swap to compare outputs without spending $$ on real API calls.
acme-q2-prod847 calls · 2.1 MBbravo-spring-baseline312 calls · 0.9 MBtone-classifier-incident47 rejected · 0.2 MBacme-q2-prod on Haiku · cost & quality deltatone-classifier-incident on rev 143 · check 47 fail-cases pass📚 Golden Cases
3 collections · 15.7K chunks indexed · 2,140 retrievals this week · $0.42 embed cost MTD
generate_promo_textYou are a copywriter for Acme Bank. Tone: professional, warm. Use insights from the Golden Cases library: [KB chunks · 3 retrieved · 1,247 tokens] ─ "When addressing B2B audiences in regulated industries…" (Brand_book_v3 p.14, 0.91) ─ "Fintech requires precision over flair…" (Tone_of_voice.md, 0.83) ─ "For senior decision-makers, lead with outcomes…" (Brand_book_v3 p.22, 0.74) Brief: Q2 launch for Acme Bank's new B2B fintech product…
kb_lookup_tonekb_lookup_brandkb_lookup_voicekb_lookup_compliance🧪 Hypothesis
2 hypotheses running · 2 designed · 1 drafting · 8 adopted / 23 total (35%)
industry=fintechpromo_copywriter@conversational- acceptance_ratio — best aligned with your hypothesis
- refusal rate — guardrail (rises if tone too informal for compliance)
- brand_safety LLM-judge — guardrail (catches off-brand outputs)
- cost/run — guardrail (informal tone often → shorter outputs → minor cost change)
- Run cost: ~$8.60 (480 runs × $0.018)
- LLM-judge eval: ~$2.40 (Haiku on 480 outputs)
- Total: ~$11 for verdict
- Compliance OK — same model, prompt change only
- Population is not balanced: Acme Q2 has 3× more traffic than Bravo Spring → consider stratified split or per-campaign analysis.
- Last similar hypothesis ("Drop tone_classifier") was rejected −9pp 5 days ago. Reviewer should check it doesn't conflict.
primary_winner: acceptance_ratio[treatment] > acceptance_ratio[control] + 0.02
AND p_value < 0.05
guardrails:
- refusal_rate[treatment] - refusal_rate[control] < 0.02
- brand_safety_judge[treatment] >= 7.0
- cost_per_run[treatment] < cost_per_run[control] * 1.20
adoption_decision: ALL_PASS → suggest adopt; ANY_FAIL → reject and document why
🤔 Allocation Adviser
Decide who should do this task: human, AI, deterministic code, or hybrid · 23 decisions logged this quarter
| Dimension | 1 | 2 | 3 | 4 | 5 | What this means |
|---|---|---|---|---|---|---|
| Volume | ○ | ○ | ○ | ○ | ~300 outputs/quarter — too much for hand-craft, too little for hardcore opt | |
| Determinism | ○ | ○ | ○ | ○ | Creative copy — many valid answers, no exact match | |
| Reversibility | ○ | ○ | ○ | ○ | Public landing page — bad output is recoverable but visible | |
| Tacit knowledge | ○ | ○ | ○ | ○ | Brand voice in KB; some judgement needed for B2B audience | |
| Time sensitivity | ○ | ○ | ○ | ○ | Daily turnaround acceptable — not real-time | |
| Eval availability | ○ | ○ | ○ | ○ | Eval suite exists (47 cases · 92% pass) — can validate AI output | |
| Liability | ○ | ○ | ○ | ○ | Brand-only liability — no regulatory or contractual claims |
| Option | Time/output | Cost/output | Quality (LLM-judge) | Throughput | Verdict fit |
|---|---|---|---|---|---|
| HUMAN_ONLY (Jess hand-crafts) | ~25 min | $8.50 | 9.4 / 10 | 3 / day | over-built |
| HUMAN_LED_AI_ASSIST (chat) | ~12 min | $4.30 | 9.2 / 10 | 5 / day | option |
| AI_LED_HUMAN_REVIEW (recommended) | ~3 min review | $0.42 | 8.9 / 10 | 30+ / day | ✓ best fit |
| AI_AUTONOMOUS (no review) | 0 | $0.42 | 8.4 / 10 ⚠ | unlimited | brand risk |
| DETERMINISTIC_NO_AI | n/a | $0 | n/a | n/a | infeasible (creative) |
td-2026-04-19-promo📝 Prompt Review
2 pending review · 5 merged this week · avg wait 4h · 100% eval-passed
tone_classifier @ stricter-rubric
in review
"Tightened the tone classification thresholds to reduce false positives. New rubric explicitly penalizes informal language for B2B contexts."
stricter-rubricheadline: string · ≤80 but eval cases all stay under 60 chars. Branch wording "return options" increases array-shape risk → would trigger TypeError: x.toLowerCase in compose_email. 3 of 100 cassette samples returned an array on this branch (0 on main).brand_voice. Acme Q2 brand-voice incidents (last week) suggest this dimension needs ≥8 cases to be representative.stricter-rubric closed · sergii notified💳 AI Usage Hub Grafana Cloud Metrics
$1,247 of $3,000 spent (42%) · forecast EoM $2,840 · advisor found $164/mo savings
| Team / cost-center | Owner | Budget | Spent MTD | Used | Forecast | Hard stop | Notify | |
|---|---|---|---|---|---|---|---|---|
| Marketing Acme client:Acme | Rebecca | $1,500 | $524 | 35% |
$1,420 | ✓ at 100% | Slack #acme | ⋯ |
| Marketing Bravo client:Bravo | Rebecca | $1,000 | $398 | 40% |
$952 | ✓ at 100% | ⋯ | |
| Internal experiments internal | Neil | $500 | $325 | 65% |
$580 | ✓ at 100% | Slack #ai-internal | ⋯ |
| Content team Charlie client:Charlie | Jess | $300 | $142 | 47% |
$305 | ⚠ at 95% (rare) | ⋯ | |
| Echo Black Friday client:Echo | Sergii | $200 | $198 | 99% ⚠ |
over cap | ❌ stops at 100% | PD + Slack | ⋯ |
| Workspace total: $1,587 spent of $3,500 in team budgets · $123 not allocated to any team | ||||||||
client:, project:, team:) automatically attribute LLM cost to the right cost center. Team owner is notified at 50/80/100% of budget. Hard-stop at 100% means new LLM calls are blocked workspace-wide for that tag (in-flight runs complete).
extract_intent from Opus to Haiku → save ~$78/mocompose_landing_html → save ~$54/mosummarize_brief with regex → save ~$32/moRuns above $50 require owner approval before start.
When enabled, your Anthropic / OpenAI keys will be used. Your invoice from the provider; we are no longer in the billing path.
🎯 Risk Classification
71 CC-items classified · 4 MULTI_LAYER · 18 HUMAN_REVIEW · 7 CO_CREATION · 42 TRUST
compose_email to TRUST · blockedTRUST_BY_DEFAULT
42 CC-items · default for most typesOutputs are reversible, low-impact, audience small or internal. Publishing is automatic. Failures don't damage brand.
| Control | TRUST | REVIEW | CO-CREATION | MULTI-LAYER |
|---|---|---|---|---|
| Eval Suite required | ≥1 case | ≥10 cases | ≥10 + LLM-judge | ≥30 + 2 judges |
| Pre-publish gate | auto-pass | block on regression | block + human OK | block + 2 humans |
| Reviewer required | — | 1 reviewer | 1 reviewer + author | 2 reviewers (incl. compliance) |
| Hallucination detector | off | warn-only | warn-only | block on suspect |
| PII redaction | passive | passive | active | active + audit log |
| Resilience policy | any | any | require exact model OK | Require Exact Model |
| Fallback to other model | allowed | allowed (logged) | warn user | blocked |
| Audit retention | 30 days | 90 days | 1 year | 7 years |
| Customer report inclusion | aggregated | per-run available | per-run + transcript | per-run + judge rationale |
| Output contract (type/shape) | recommended | required · warn-only | required · block + retry | required · block + halt + audit |
| Artifact immutability after approval | none · agent free-edit | seal-on-merge · agent edit allowed but flagged in audit | seal-on-merge · agent edit auto-breaks seal → re-review required | seal-on-merge · any edit (human or agent) requires Change Plan + 2nd reviewer |
| Intent allowlist (Guardrails) | off | allowlist + warn | allowlist + refuse | allowlist + refuse + audit |
| Context grounding floor | — | warn if <0.6 | require ≥0.7 or refresh KB | block if <0.8 |
| Vector-store availability (Ragie) | graceful · skip retrieval, log warning | graceful · alert · run completes without RAG context | block if RAG unavailable · queue + retry | block + halt + page on-call · no degraded retrieval |
| Eval coverage floor | — | ≥60% | ≥75% | ≥90% across all dimensions |
| Reviewer diligence friction | — | 2/4 mandatory checks | 3/4 mandatory + AI second-opinion | 4/4 + AI second-opinion + adversarial cassette |
| Reviewer reliability floor | — | — | ≥0.65 or 2nd reviewer auto-added | ≥0.80 or compliance reviewer added |
| CC-item | Type | Campaign | Auto-classified by | Last edit | |
|---|---|---|---|---|---|
extract_intent | json_prompt | Acme Q2 | type-default | 2h ago | › |
summarize_brief | prompt | Acme Q2 · Bravo | type-default | 1d ago | › |
tone_classifier | json_prompt | shared (3) | type-default | 14:10 today (Jess) | › |
fetch_brief | http | Acme Q2 | type-default | 3d ago | › |
post_process | function | Acme Q2 | type-default | 5d ago | › |
| + 37 more | |||||
<eval> consumers.🕸 Impact Graph
Selected: tone_classifier · 17 downstream artifacts · 3 will break · 4 campaigns · 8 eval cases · 2 contracts · 1 MCT
promo.schema.jsontone_result.contractretrieve_brand_voiceretrieve_compliancetone_classifier would break 3 downstream consumersexperiment-tone-v2. Impact analysis shows: 1 contract violation (tone_result.contract shape would change), 1 MCT (mct_b2b_finance replicates to 12 client tenants), 1 stale eval suite (tone_judge_v2 uses brand_guides rev 7, source updated 6w ago). Open as Change Plan to capture mitigation steps before merge.- •
tone_result.contract— output shape change - •
mct_b2b_finance— 12 client tenants downstream - •
compose_email—.toLowerCase()on changed shape
- • 2 eval suites · auto re-run on merge
- • 2 PRs in queue · new diligence required
- • 4 campaigns · expect acceptance drift 5-15pp
- •
chat_assistant_dag - • Quality Gate suite · normal trigger
- • Customer Feedback — pattern detector ready
- • Weekly Review · cards refresh on Sun
📋 Change Plan
3 active · 1 awaiting approval · 12 merged this week · avg plan→merge cycle 2.4d
CP-104 · tone tightening
awaiting approvalAuthor jess@workspace · Created 38m ago · Reviewer neil (auto-routed: tone_classifier owner)
- Coordinate with neil (mct_b2b_finance owner) — schedule MCT replication for off-peak window
- Add Output Contract migration note:
tone_result.contractv2 (backwards-compatible coercion for 14 days) - Re-fresh
brand_guidesKB chunk before publish (currently stale 6w) - Notify Rebecca (Acme account) of expected acceptance drift 5-15pp on Day 1
- • Acceptance ratio recovers to ≥85% within 2h post-merge
- • Zero new
OUTPUT_CONTRACT_VIOLATIONevents in 24h - • Acme CFO confirms tone fix on next-day campaign
tone_result.contract shape, but doesn't list the contract migration deliverable explicitly. Recommend adding "+ Bump contract version to v2" to Scope.compose_email author. Recommend: stick with current approach but document the trade-off here.📅 Weekly AI Quality Review Grafana Cloud Reporting
This Monday Apr 19, 10:00 UTC · 7-card agenda · 4 attendees · 4 action items pending
| Prompt | Direction | Δ | Owner | Action |
|---|---|---|---|---|
tone_classifier | ↓ acceptance | −17pp (resolved) | Jess | 5 eval cases added |
json_extractor | ↑ refusal rate | +5× over week | Sergii | investigate provider policy update |
compose_email | ↑ length | +18% avg tokens | Jess | tighten output spec |
| Action | Owner | Due | Linked to | |
|---|---|---|---|---|
| Disable direct edits on shared prompts (require branch) | Neil | Apr 24 | Card 1 | |
| Investigate json_extractor refusal spike | Sergii | Apr 22 | Card 2 | |
| Decide on Hypothesis #4 adoption for Charlie | Rebecca | Apr 22 | Card 4 | |
| Schedule compose_landing_html cache rollout | Neil | Apr 26 | Card 5 |
Every Monday 09:30 UTC the system auto-builds this deck from:
- Quality Insights · alerts & drift
- Reviews · merged + bypassed
- Hypothesis · running & adopted
- Usage Advisor · realized + queued
- KB · updates & impact
- Tickets & action items closed since last week
🔀 Business flows
13 end-to-end flows · sequence diagrams of who triggers what, who responds, where decisions happen
🗺 Feature map
19 features grouped by lifecycle stage · for the demo conversation with the customer
TypeError: x.toLowerCase in <eval> nodes before MCT-key bugs reach prod.- Campaign Explorer → "Acme Q2" pinned card → recommended action banner.
- Quality Insights → 17pp drop · likely cause (prompt edit) · Hallucination Detector card.
- Customer Feedback → Acme CFO complaint matches the pattern — "3rd this week".
- Inspector → run #4828 with the actual bad output.
- Workflow Editor → see the prompt's position in the DAG · linter · cost preview.
- Prompt Studio → fix → Publish → Quality Gate blocks regression with Cassette replay validation.
- Prompt Review → teammate change with merge preview.
- Always-On AI → policy editor + simulator.
- Model Catalog → BYOK + Opus sunset planning.
- Golden Cases → Test query for retrieval.
- Hypothesis → "tone → conversational" interim verdict.
- Usage → Team budgets + Advisor $164/mo · "Apply with eval check".
🏷 Vendor landscape — what we'd build vs. what we'd buy
9 categories of AI infrastructure · 22 vendors surveyed · 5 marked as drop-in for our prototype today
- • Ragie — RAG / vector store (Chris's branch)
- • Langfuse → Inspector
- • Braintrust → Quality Gate
- • Portkey → Always-On AI
- • PromptLayer → Prompt Studio
- • Output Contract / Intent Validation
- • Hypothesis Lab
- • Usage Hub (cost optimization layer)
- • AI Second Opinion / Coverage Report
- • Impact Graph / Change Plan / Artifact Seal
- • Risk Classification matrix
- • Quality Insights — structured-logging dashboard at
veprod.grafana.net
- • Inspector → Cloud Traces (Tempo, OTel LLM)
- • Always-On AI → Cloud SLO · OnCall · Incident
- • Usage Hub → Cloud Metrics + anomalies
- • Customer Feedback → NPS/CSAT panels
- • Activity timeline → Cloud Logs (Loki)
- • Weekly Review → Cloud Reporting (PDF export)
- • Source-of-truth governance: Golden Cases collections, freshness-vs-source tracking, KB drift monitoring.
- • Grounding score: 6-signal cross-check (KB freshness, retrieval relevance, drift, memory loss, budget, var resolution).
- • Stale-KB consumer alerts: "38 prompts depend on a chunk that hasn't been re-indexed in 6 weeks".
- • Workflow integration: retrieval as first-class CC-item with type-checked I/O.
- • Audit grade: per-run RAG_QUERY events with chunk-level provenance for compliance.
RAG-ccItem.- • Business-context tagging: campaign ID, asset type, MCT lineage. Vendor sees calls; we see which campaign for which client.
- • Domain events: SEAL_VERIFIED, OUTPUT_CONTRACT_OK, INTENT_VALIDATED, RAG_UPLOAD_FAILED, ARTIFACT_SEAL_BROKEN. These are our policy primitives, not observability.
- • "Reproduce in <5 min" UX: Inspector → Studio → Eval Gate → Customer Feedback links — that's a product, not a trace viewer.
- • Audit-grade retention: 3-tier retention with PII redaction tied to risk classification.
- • Campaign-quality rubric: brand voice, promo correctness, regulated-industry tone — vendors don't know what "good copy for Acme Bank" looks like.
- • Coverage Report: 6-dimension matrix that tells where the suite is weak, not just which cases pass.
- • Auto-capture from Inspector + Customer Feedback: the loop from real complaint → permanent regression test.
- • Publish-gate enforcement: the gate sits in the Studio publish flow with seal-and-revert mechanics. Vendor would be the eval engine, we own the policy.
- • Risk tab: Output Contract, Intent Validation, runtime policy declared per CC-item.
- • Change-Plan-gated edits: can't open Studio for production prompts without an approved plan in scope.
- • Impact Graph reverse-deps: "this edit will break 12 MCT tenants downstream" — vendor doesn't see your DAG.
- • Linter with domain rules: contract violation risk, role-tagging on user data, KB usage hygiene.
- • Multi-stakeholder review workflow: AI Second Opinion + diligence checklist + reviewer reliability score.
- • Per-CC-item policy: "MULTI_LAYER prompts must use claude-opus-4.7, no fallback"; vendor handles routing, we set the policy.
- • Tier-aware enforcement: Risk Classification matrix decides what fallback is allowed for each tier.
- • Compliance guards: Require Exact Model + region restrictions tied to client contracts.
- • Reliability event correlation: fallback events surface in Inspector with cost overhead tagged.
- • JSON schema validation is commodity (Pydantic / Instructor cover it natively).
- • Output Contract is more than schema — it's a 4-phase lifecycle (write / review / publish / runtime) tied to risk tier and downstream consumer analysis. Vendors don't model that.
- • Intent Validation is a cheap-classifier-as-gate pattern that lives before the LLM call; vendors validate after.
- • Domain rules (no competitor names, no hallucinated SKUs) require workspace-specific policies vendors don't host.
- • No vendor is purpose-built for "LLM Hypothesis Lab" — they're general A/B platforms.
- • PICO-style framing + AI-designed experiment + pre-registered analysis is our authoring layer; vendor would only run the math.
- • Cassette-replay-based experiments (no real prod cost) is unique to our stack — vendors don't model it.
- • Adopted-hypothesis → Change Plan transition is workflow logic, not statistics.
- • No "Datadog for LLM cost" winner exists yet; everyone builds on top of Helicone/Langfuse primitives.
- • Usage Advisor ("switch this prompt to Haiku → save $78/mo") is a domain-aware optimizer — vendors give numbers, we give recommendations.
- • Per-team budgets + chargeback PDF + BYOK live in our workspace model; vendors are blind to it.
- • Cost-bounded retries tied to risk tier — that's gateway + policy, not a cost dashboard.
- • Maturity is real for security red-teaming; we're focused on quality coverage.
- • AI Second Opinion is workflow-aware (cassette replay + coverage report + contract risk) — vendor red-teams in isolation.
- • Coverage Report identifies which dimension is undertested; vendors generate cases without that targeting.
- • Promptfoo is interesting as a reference for case-generation patterns; not a drop-in.
📖 Documentation
Product Guide · auto-generated from the codebase by deep-code-analysis