Campaign Explorer
Acme Marketing · Apr 21
⚠ Recommended · 1 action
Acceptance ▼ 17pp on Acme Q2 since 14:10 · auto-correlated with tone_classifier rev 142 edit. Eval suite not run.
87.3%
Acceptance
▲ 2.1pp
0.19%
Hallucinate
blocked
$1.2k
7d cost
▼ 8% lat
⭐ Pinned campaigns · 3
Acme Q2 Launch
Education · H1-2026 · needs attention
Bravo Spring Promo
Retail · 94% acceptance · healthy
Golf B2B Outreach
Running now · 3 assets in review
All campaigns · 7 tap to select
Acme Q2 Launch81% ▼
Bravo Spring94%
Charlie Holiday91%
Delta Newsletter96%
Echo Black Friday78%
Foxtrot Loyalty93%
Golf B2B Outreachrunning
Acme Q2 Launch Education · H1-2026
needs attention 47 runs today 4 prompts 2 MCT tenants
Acceptance
81%
▼ 17pp
7d cost
$248
▲ 12%
CSAT
4.1
▼ 0.3
Contract ✓
96%
2 vio
17pp drop since 14:10 after tone_classifier rev 142. Eval suite not run. Acme CFO complaint referenced output.
📅 Activity · today Grafana Cloud Logs
14:12
Quality alert · Acme Q2 ▼17pp
auto-opened via Inspector
13:40
Hypothesis running · tone-b2b
day 2/5 · projected ADOPT
11:02
Bravo Spring · 38 outputs shipped
0 contract violations
09:15
Acme CFO · 👎 captured as regression
"AUM figure made up"
📡 Monitor · production health
🔍 Inspector Every run · prompt · events · view only
alert 📈 Quality Insights Trends · hallucination · grounding
new 🕸 Impact Graph 3-hop blast radius
new 💬 Customer Feedback 👍/👎 · triage · patterns
💳 Usage Hub Budget · chargeback · savings
📅 Weekly Review 7 auto-prep cards
🏗 Inspect · authored artifacts
new 🔀 Workflow Editor DAG · linter findings · view
✏️ Prompt Studio Current draft · diff · revisions
3 🚦 Quality Gate Publish checks · regressions
2 📝 Prompt Review Pending PRs · diligence
🎯 Risk Tiers Controls matrix · compliance
new 📋 Change Plan Active plan · sealed artifacts
🧪 Decide · ask the lab
🧪 Hypothesis Frame a PICO · run it
new 🤔 Allocation Adviser Score 7 dims · get verdict
⚙ Settings · read-only view
Always-On AI Fallback chain · policy
new 🧠 Model Catalog Pricing · sunsets · BYOK
🗺 Feature map All 19 features · grouped
🏷 Vendor landscape Buy vs build · 9 categories
🔀 Workflow Editor

DAG editor with type-checked I/O

promo_dag · 18 nodes · 24 edges · 2 linter warnings

⚠ Cannot publish yet
generate_hero (MULTI_LAYER) can publish without human approval. Fixing the graph happens on desktop.
Canvas · selected path tap a node
startSTART
httpfetch_brief120ms
kb_lookuptone_chunks1.2s
json_promptextract_intent1.4s · $0.04
promptapply_tone1.8s
dallegenerate_heroMULTI_LAYER
human_approval⚠ missing in graph
assetpublishstaging
⚠ Linter · 2 findings
generate_hero
no human_approval before publish · critical
extract_intent
outputs audience, downstream unused · unused output
Selected node · extract_intent
json_prompt HUMAN_REVIEW
~$0.04/run · 1.4s avg · 1 retry typical
Inputs
brief, kb_results
Outputs
intent, audience
Run estimate · whole graph
Cost
$1.20
per run
Latency
12.4s
p50
Tokens
42.1k
avg
LLM calls
24
per run
✏️ Prompt Studio

promo_copywriter rev 142 draft

On branch experiment-tone-v2 · 2 linter warnings · seal 🔓 broken

📐 Output contract is set
JSON promo.schema.json · enforced at linter → review → publish → runtime. Catches .toLowerCase() bugs before MCT prod.
Edit
Diff vs main
Variables
Risk
System prompt · 847 chars · read-only
You are promo_copywriter. Given a brief and tone, produce a JSON object: { headline, bullets[3], cta }.

Strict rubric: text MUST contain at least 3 indicators of the chosen tone (e.g. for "professional": industry vocabulary, complete sentences, no contractions).

If unsure, return "uncertain" — do NOT guess.
−6 Choose: professional / casual / urgent / friendly
+6 Choose: professional / casual / urgent / friendly
+7 Strict rubric: text MUST contain at least 3 indicators
+8 of the chosen tone (e.g. for "professional": industry
+9 vocabulary, complete sentences, no contractions).
+10 If unsure, return "uncertain" — do NOT guess.
{{brand}}"Acme Bank"
{{audience}}"B2B CFO"
{{tone}}"professional"
{{kb_results}}3 chunks
Tier
MULTI_LAYER
Contract
JSON · block
Reviewers
2 required
Audit
7 years
  • Halluciation detector: block on suspect
  • Fallback to other model: blocked
Cost preview · per 1 call
claude-3-5-sonnet$0.018 · 824ms
claude-3-haiku$0.003 · 410ms
claude-3-opus$0.082 · 2.1s
gpt-4o$0.015 · 780ms
gemini-1.5-pro$0.008 · 620ms
⚠ Linter · 2 warnings
Rubric wording change unverified
diff lines 7–10 · no eval case covers "uncertain" refusal path
Contract unchanged but meaning drifted
bullets cap = 3 still · AI Second Opinion flagged tonal shift
Used by · 4 CC-items
Acme Q2 Launch96%
Bravo Spring Promo94%
Echo Black Friday78%
+ 1 more
Memory · prior decisions
  • 2026-03-04 · rejected "use 5 bullets" — downstream email wraps at 3
  • 2026-02-18 · adopted "JSON-only output" — closed a 6-day MCT incident
  • 2026-01-22 · rejected switch to gpt-4o — cost +35%, acceptance flat
Revisions
rev 142 · draftJess · 1h
rev 141 · mainRebecca · Mon
rev 140Neil · last week
rev 139Jess · 2w
🚦 Quality Gate

Suite: promo_copy_quality

47 cases · against rev 141 (main) · 3 regressions

Improved
6 ▲
Regressed
3 ⚠
Unchanged
38
Cost Δ
+12%
Criteria · 5 checks
OUTPUT_CONTRACTw52/47 ✗
CONTAINSw1
MAX_LENGTHw1
JSON_SCHEMAw2
LLM_JUDGEw33/10 ✗
Regressions · 3
json-strict-schema
✗ JSON_SCHEMA · missing field "cta"
short-tweet-280
✗ MAX_LENGTH · 312 > 280
tone-rubric-judge
✗ 3/10 · below 0.7 threshold
Coverage by dimension
tone variations4/5
locale & language2/3
edge lengths3/4
adversarial input1/4
JSON schema variants3/3
📼 Cassette & Replay
47 real production calls captured · replayable against any draft on desktop · zero API spend when re-running.
📝 Prompt Review

2 pending · 5 merged in 7d

GitHub-style PRs for prompt changes · diff + eval + merge preview + rollback arm.

Pending · 2
tone_classifier · rev 142
Jess · 1h · 3 regressions 2 reviewers req
compose_email · rev 28
Sergii · 3h · all pass
Before you approve · mandatory
Eval suite green
No regressions3 ✗
Contract unchanged
AI second opinionflagged
Rollback condition set
Merge preview · impact
On merge:
  • Swap tone_classifier rev 141 → 142
  • Touch 4 campaigns (Acme Q2, Bravo Spring, Charlie Holiday, Golf)
  • Arm rollback if acceptance < 80% within 2h
  • Notify 4 affected reviewers
Reviewer reliability · 30d
Rebecca42 reviews · 0.89
Neil31 reviews · 0.82
Jess18 reviews · 0.64
Below 0.65 → workspace policy adds a 2nd reviewer automatically.
Approving / requesting changes happens on desktop. Mobile shows the state of the queue so you know what's waiting.
🎯 Risk Classification

4 tiers · controls matrix

Tier drives eval coverage, reviewers, retention, fallback policy. Misclassification attempts are blocked & logged.

Tiers & items
TRUSTTrust by default16 items
REVIEWHuman review24 items
CO-CRECo-creation14 items
MULTIMulti-layer8 items
Controls matrix · REVIEW vs MULTI
ControlREVIEWMULTI
Eval cases≥10≥30 · 2 judges
Reviewers12 + compliance
Pre-publish gateblock regressionblock + 2 humans
Halluciation det.warn-onlyblock suspect
PII redactionpassiveactive + audit
Fallback modelallowed · loggedblocked
Audit retention90 days7 years
Output contractwarn-onlyblock + halt
Intent allowlistallowlist + warnallowlist + refuse
Misclassification log · 30d
generate_hero → TRUST
blocked · dalle always MULTI
summarizer → REVIEW
blocked · MCT needs MULTI
Compliance posture
SOC2 · ready GDPR · ready ISO 42001 · partial EU AI Act · ready
📋 Change Plan

CP-101 · intent allowlist approved

Plan sealed by Neil · 2d ago · in scope: promo_copywriter, compose_email

Plan · 4 lines
Scope
tone_classifier rev 141→142 on branch experiment-tone-v2
Contract impact
tone_result.contract: string{value,confidence} · 5 consumers
Rollout
Cassette replay · staged 10% → 50% → 100% over 3h
Rollback
Auto-revert if acceptance < 80% within 2h
Seal status · 2 artifacts
promo_copywriter
🔒 sealed · Neil · hash ac49…
tone_classifier
🔓 broken · agent edit 14:23
Affected reviewers · 4 notified
Rebecca Neil Sergii Jess (author)
🧪 Hypothesis

Frame a new hypothesis

Fill the PICO boxes · the Lab calculates sample size and designs the experiment.

PICO canvas required
Guardrail metrics must not regress
Effect size & window
🎛 Auto-designed by the Lab
Sample / arm
480
power 0.8
Estimated cost
$34
both arms
Stop rule
CI excl. 0 · 24h
Runs on
cassette + live
Pre-registered analysis stored when the hypothesis is launched · prevents late-metric-switching.
AI critique of your draft
Ready to launch
Population is large enough for the MDE you picked. Primary metric is well-defined. No overlapping hypothesis on the same prompt.
Projected impact · if adopted
Acme Q2 Launch+2.8pp · ~$62/mo saved
Bravo Spring+1.9pp · flat cost
Charlie Holiday+0.8pp · +$9/mo
Projections from cassette replay against the selected population.
Pipeline · this workspace
▶ runningconversational-tone-b2bDay 2/5
▶ runninghaiku-migration-intentsDay 4/7
📋 designedmemory-bank-retrievalidle
✓ adoptedprompt-caching-phase1−18% cost
✗ rejectedgpt4o-for-creativeno gain
🤔 Allocation Adviser

Score the task, get the verdict

Tap the dots to rate each dimension 1–5. The verdict updates live.

7 dimensions · tap a dot 1–5
Volumetap to see why
Determinismtap to see why
Reversibilitytap to see why
Tacit knowledgetap to see why
Time-sensitivitytap to see why
Eval availabilitytap to see why
Liabilitytap to see why
AI verdict
AI_LED + HUMAN_REVIEW
High volume + high liability + reversible → AI handles first pass, human reviews on refusal / confidence < 0.7. Blocks both default-to-AI and default-to-human.
HUMAN_ONLY HUMAN_LED_AI AI_LED_REVIEW ✓ AI_AUTONOMOUS DETERMINISTIC
Estimates · per month
Throughput
~2,400
decisions
AI cost
$86
@ Haiku
Review load
~240
10% refer
Risk tier
CO-CRE
auto-set
Compare with other verdicts
VerdictCost/moReviewRisk
HUMAN_ONLY$0low
HUMAN_LED_AI$24~1,680low
AI_LED_REVIEW$86~240med
AI_AUTONOMOUS$580high
DETERMINISTIC$00
Linked artifacts · will be created
promptintent_classifiernew · tier CO-CRE
eval suiteintent_quality10 seed cases
review policyrefer-if-conf < 0.7auto-wired
Decision history · this workspace
AI_LED_REVIEWpromo_copywriterNeil · Apr 12
HUMAN_LED_AIbrand_safety_reviewRebecca · Mar 30
DETERMINISTICcurrency_formatSergii · Mar 22
HUMAN_ONLYcompliance_attestLegal · Mar 11
AI_AUTONOMOUStag_extractionChris · Feb 28
Workspace verdict distribution · this quarter
AI_LED_REVIEW████████░░ 38%
HUMAN_LED_AI████░░░░░░ 22%
DETERMINISTIC███░░░░░░░ 18%
AI_AUTONOMOUS██░░░░░░░░ 14%
HUMAN_ONLY██░░░░░░░░ 8%
Skew shows workspace bias · too much AI or too much human both get flagged.
🔍 Inspector

run #4823 · generate_promo_assets

success · 14:32 · 12.4s · 24 LLM calls · $1.20

Grafana Cloud Traces
Steps
Input
Resolved
Output
Events
fetch_brief
http · 120ms
kb_lookup_tone
Ragie · 3 chunks · top 0.91 · 1.2s
extract_intent
json_prompt · ↻×1 · 1.4s · $0.04
apply_tone
prompt · ⇄ openai fallback · 1.8s · $0.06
generate_promo_text
prompt · 4.1s · $0.42 · selected
post_process
function · 40ms
generate_dalle_hero
dalle · 3.2s · $0.04
moderate_image
guard · 200ms
compose_email
prompt · 1.1s · $0.18
compose_landing_html
prompt · 1.6s · $0.22
review
asset · human · 2.3s
publish
asset · skipped (review-only)
Step inputs · resolved
{{brief}}"Q2 treasury positioning…"
{{tone}}"professional"
{{audience}}"B2B CFO"
{{brand}}"Acme Bank"
KB chunks · brand_guides_acme · rev 7 · top-K 3 · reranker on
Vars substituted · KB chunks expanded
brand: "Acme Bank"
audience: "B2B CFO"
tone: "professional"
kb:
  - "Q2 offerings focus on treasury"
  - "Compliance-first positioning"
  - "Onboarding 48h SLA"

Produce promo JSON { headline, bullets[3], cta }.
Validated against promo.schema.json
{
  "headline": "Q2 brings precision to your treasury.",
  "bullets": [
    "Real-time cash visibility for CFO teams.",
    "Built for regulated industries — SOC2, HIPAA-ready.",
    "Onboard in days, not quarters."
  ],
  "cta": "Book a 20-min walkthrough"
}
01.155BUDGET_CHECK_PASSED · est $0.018
01.487RAG_QUERY_OK · 3 chunks 0.91
01.890LLM_CALL_OK · sonnet 824ms
01.905OUTPUT_CONTRACT_OK
02.100SEAL_VERIFIED
05.214STEP_COMPLETED
LLM call · generate_promo_text · attempt 1/3
Model
sonnet-3.5
Tokens
4.1k in · 312 out
Latency
824ms
Cost
$0.018
Call id 8f3a2b… · valid JSON ✓ · contract matched
Retrieval · KB provenance
brand_guides_acme · rev 7
chunk 14/180 · score 0.91
compliance_fintech
chunk 3/52 · score 0.87
tone_of_voice_v3
chunk 8/64 · score 0.71 · ⚠ stale 6w
Grounding signals · this run
KB freshness0.94 ok
Retrieval quality0.92 ok
Drift (24h)0.78 watch
Memory loss0.96 ok
Budget fit0.91 ok
Var resolution0.82 watch
Tags on this run
client:Acme project:Q2 priority:high risk:MULTI_LAYER mct:b2b_finance
Run list · recent tap to switch
#4823 · generate_promo14:32 · ok
#4828 · Bravo Spring11:02 · ok
#4819 · Delta Newsletter10:14 · ok
#4811 · Charlie Holiday09:55 · ok
#4712 · MCT contract ✗09:48 · contract
#4708 · seal broken09:32 · seal
#4694 · rag_upload fail08:51 · rag
Workspace filters
📈 Quality Insights

Workspace · 7 days

Acceptance · regressions · hallucinations · grounding

Grafana Cloud · live
⚠ Acme Q2 · acceptance ▼ 17pp
Drop since 14:10 after tone_classifier edit. Eval suite was not run. Likely cause: skipped rubric.
Acceptance
87.3%
▲ 2.1pp
Schema violations
1.2%
▲ 0.4pp
Refusals
0.4%
▲ 5×
NPS (4w)
47
▲ 4
Acceptance · 14-day trend
2w ago▼ today
Acceptance by campaign
Delta Newsletter96%
Acme Q294%
Foxtrot Loyalty93%
Bravo Spring91%
Charlie Holiday89%
Echo Black Friday78% ▼
Delta A/B Test62% ▼
🧠 Hallucination Detector · 7d Grafana Cloud alerting
Scanned
2,140
Suspect
23
1.07%
Confirmed
4
all blocked
Grounding
0.91
≥ 0.8
🧭 Context Grounding Monitor
KB freshness0.94
Retrieval quality0.92
Drift (24h)0.78
Memory loss0.96
Budget0.91
Var resolution0.82
4.2× correlation with hallucinations · predicts before they ship.
🕸 Impact Graph

Blast radius · tone_classifier

Editing rev 142 would touch 17 artifacts across 3 hops.

🚫 Pending edit breaks 3 downstream
1 contract violation · 1 MCT (12 tenants) · 1 stale eval suite. Mitigation is captured in the associated Change Plan.
HOP 1 · direct consumers · 6
tone_result.contractwill break
mct_b2b_finance12 tenants
tone_judge eval suitere-run
promo_dag · node 4re-review
brand_voice eval (4)cov ▼ 65%
chat_assistant_dagunaffected
HOP 2 · transitive · 5
<eval> compose_email.toLowerCase risk
12 client tenantsMCT breaks
Reviews queue2 PRs
Quality Gate suiteauto-runs
hypothesis · tone-b2bsnapshot
HOP 3 · surface · 6
3 client micrositesrecoverable
Weekly Review deckcard 2 edit
customer_feedback_aggpattern refresh
SOC2 audit bundleno change
NPS 4w · Acmetracked
Monthly PDF · Acmeregen
🔎 Killer query · ask the graph
Saves as workspace monitor · alerts when the answer grows.
Saved queries
Every prompt with stale KB7 · ▲ 2
Contracts w/ >5 consumers3
Nodes w/o eval coverage4 · ⚠
MULTI_LAYER without human_approval1 · block
Selected artifact
prompt tone_classifier
rev 142 → 143 (draft) · author Jess · branch experiment-tone-v2
💬 Customer Feedback

End-customer 👍 / 👎

NPS 47 · CSAT 4.3/5 · 82% 👍 ratio ▼ 3pp

Grafana Cloud Metrics
🔍 Pattern detected
6 complaints in 10 days about "tone too formal for B2B fintech". Open as hypothesis to prove the fix works.
Triage queue · 8
Acme CFO · 👎
"AUM figure made up. Not in report."
Bravo PM · 👎
"Tone too formal for tech audience"
Charlie · 👎
"Copy repeats bullet 2 in bullet 3"
Delta · 👍
"Nailed the compliance framing"
Selected · Acme CFO
"The promo claims Acme holds $2.3B AUM — this is wrong and nowhere in the report we sent."
hallucination contract violation run #4823
Campaign ratings · 4w
Campaign👍👎%
Delta Newsletter42295%
Bravo Spring38588%
Acme Q2521281% ▼
Charlie Holiday28390%
Echo Black Friday21872% ▼
💳 Usage Hub

$1,247 / $3,000 · 41%

Workspace budget · 18 days remaining · advisor has 3 savings

Grafana Cloud Metrics
Monthly spend · daily
Day 1Spike · Apr 15Today
💡 Advisor · 3 savings found
Switching 3 prompts to cheaper models saves $127/mo. Quality on cassette = equivalent or better.
Savings · cheap-model swaps
extract_intent · Opus → Haiku
cassette: +1pp acceptance
−$78
post_process · Sonnet → Haiku
cassette: equivalent
−$41
compose_subject · cache 74%
prompt caching tune
−$8
Per-team budgets
Growth Ops$520 / $1,200 · 43%
Creative$482 / $900 · 54%
R&D$245 / $900 · 27%
Soft warn 50/80% · hard stop 100% (paused runs need explicit unpause).
BYOK · 2 of 4 providers
Anthropic · customer keyactive
OpenAI · customer keyactive
Google AI · platform
Bedrock · platform
📅 Weekly Review

Week 17 · Mon Apr 21

Auto-prepared Sun 22:00 · 7 cards · 4 action items

Grafana Cloud Reporting
Card 1 · Last week's actions
Roll back tone_classifier 141done
Opus → Haiku for classifierdone
Acme CFO replyin flight
Ragie partition hygieneoverdue
Card 2 · Quality alerts
  • Acme Q2 · −17pp after prompt edit · open
  • Echo Black Friday · acceptance 78%
  • Hallucination rate 0.19% · all blocked ✓
Card 3 · Customer feedback
Pattern: "tone too formal" · 6 complaints in 10d. NPS 47 (▲4). CSAT 4.3/5.
Card 4 · Hypothesis status
2 running 2 designed 2 adopted 1 rejected
Card 5 · Usage · savings this week
Advisor captured $78/mo on classifier · $41/mo on post_process.
Card 6 · GC updates
1 brand guide updated (Tone_of_voice_v4) · 8 prompts auto-picked up new rev.
Card 7 · New action items
Run tone-b2b hypothesis
Rebecca · Fri
Ragie partition audit
Chris · Wed
Acme CFO follow-up reply
Neil · today
⚡ Always-On AI

Fallback chain across providers

Anthropic → OpenAI → Bedrock · policy editor + simulator · cassette drills

Grafana Cloud SLO · OnCall · Incident
Provider status · live
Anthropichealthy 3 models
OpenAIhealthy 3 models
Google AIdegraded 2 models
Bedrockhealthy 1 model
Default fallback chain
primaryAnthropic · sonnet-3.5
↓ on 5xx × 3
fallback 1OpenAI · gpt-4o
↓ on 5xx × 3
fallback 2Bedrock · claude
gracefulcached answer · mark retry
Per-tier policy
TierRetryFallback
TRUST×3allowed
REVIEW×3allowed · logged
CO-CRE×2warn user
MULTI×1blocked
🚨 Simulator · last drills
Anthropic 5xx12:42 · 1.2s fail-over
OpenAI quotaWed · Bedrock took over
All providers downMon · graceful-degrade fired
Simulator runs against cassette · zero production impact · triggered from desktop.
Recent drills · 7d
Anthropic 5xx burst12:42 · ✓ 1.2s
OpenAI quotaWed · ✓ Bedrock
Google AI degradedtoday · active
🧠 Model Catalog

Unified across providers

Pricing · latency · context · sunset · BYOK · side-by-side compare

⚠ Opus sunset Jul 31
claude-3-opus-20240229 deprecating · 2 prompts still use it. A migration hypothesis is auto-generated and tracked in the Lab.
Anthropic · 3 models
claude-3-5-sonnet
$3/$15 · 824ms · 200K
primary
claude-3-haiku
$0.25/$1.25 · 410ms · 200K
fast
claude-3-opus
$15/$75 · 2.1s · 200K
sunset Jul 31
OpenAI · 3 models
gpt-4o
$5/$15 · 780ms · 128K
fallback
gpt-4o-mini
$0.15/$0.6 · 420ms · 128K
fast
gpt-3.5-turbo
$0.5/$1.5 · 380ms · 16K
legacy
Side-by-side · run same prompt
MetricSonnetGPT-4o
Acceptance94%91%
Cost /100$1.80$1.50
Latency p50824ms780ms
Contract OK98%93%
BYOK · 2 of 4
Anthropic & OpenAI on customer keys · we're out of the billing path · client data stays in their compliance scope.
🗺 Feature map

All 19 features by lifecycle stage

Hub · Build · Lab · Operate · Settings. Tap any card to jump into its screen.

Every feature here answers three questions: why it exists (the business problem), concrete result after shipping, and key arguments for the customer conversation. Ordered by where it sits in the workflow lifecycle.
🧭 Hub
🧭 Campaign Explorer HUB
Daily standup · every campaign, health, cost, recent activity in one place.
Problem: 5 dashboards before knowing where to focus.
Result: 30-min "where to spend my day" picture · recommended-action banner.
Args: single pane of glass · auto-correlated alerts · activity timeline aggregates from every other feature.
🏗 Build · author, test, review & ship
🔀 Workflow Editor new
Visual DAG of CC-items · the no-code pipeline builder customers actually use.
Problem: graph errors at runtime · non-engineers can't safely wire prompts → KBs → guards.
Result: shippable DAG · type-checked I/O · cost + latency known before first run · linter blocks dangerous topologies.
Args: 14+ CC-types · per-node deep-link to Studio / Eval / Risk · cycle & MULTI_LAYER-without-human_approval linter · graph-level versioning.
✏️ Prompt Studio ↪ PromptLayer
IDE for prompts · branches · A/B · cost preview · Memory tab · Risk tab.
Problem: "edit prod and see what happens" · non-engineers can't author safely.
Result: tested-before-shipping prompt · cost + latency known · reviewer signs off · 1-click rollback.
Args: 5-model cost preview · variable chips · linter · diff vs main · publish triggers Quality Gate · Allocation Adviser inline.
📐 Output Contract: declared shape (string · JSON schema · array) enforced write → review → publish → runtime.
📝 Prompt Review
GitHub-style PRs for prompt changes · diff + eval + merge preview.
Problem: Slack approvals · evidence scattered · no compliance trail.
Result: single approval queue with diff + eval + cost/latency delta · auto-rollback armed.
Args: red/green inline diff · embedded eval · merge preview lists affected campaigns · AI Second Opinion + diligence checklist · per-reviewer reliability score.
📚 Golden Cases · KB ↪ Ragie
Upload brand guides once · RAG retrieval into kb_lookup nodes · single source of truth.
Problem: brand guides copy-pasted into every prompt → bill shock, inconsistency.
Result: update one PDF → 40 campaigns pick it up · ~47% input-token reduction.
Args: drag-drop upload · 3 chunking strategies · reranker + hybrid search · test-query playground · audit retrieval · feeds Usage Advisor savings.
🎯 Risk Classification
4 tiers drive controls automatically — eval, reviewers, retention, fallback.
Problem: DALL·E to public S3 gets same thin pipeline as an internal classifier.
Result: brand-risky outputs cannot ship without human approval · compliance answer in 1 minute.
Args: 4 fixed tiers · 9×4 controls matrix · misclassification audit · SOC2 / GDPR / ISO 42001 / EU AI Act posture · enforced by Quality Gate, Reviews, Reliability.
🛡 Intent Validation + 🔒 Artifact seals: per-CC allowlist blocks out-of-scope queries before the expensive LLM call · approved artifacts hash-sealed; any edit re-triggers tier-appropriate review.
🚦 Quality Gate ↪ Braintrust
Test suites for prompts · block regressions · auto-capture bad prod cases · cassette replay.
Problem: "tested on one example, looked fine" → silent regression.
Result: growing safety net · pre-publish dialog explains failures · bypass requires written reason.
Args: 6 criterion types incl. LLM_JUDGE with calibration · cassette replay (no API spend) · auto-capture from Inspector & Feedback · Coverage Report with "70% pass on 4 cases" anti-pattern detector.
📋 Change Plan new
Review the intent to change before any artifact is generated.
Problem: reviewers approve 200-line diffs on trust · architectural debate arrives too late.
Result: motivation / scope / impact / rollback discussed at 4-line stage · plan approval unlocks Studio for listed artifacts only · audit-grade signature.
Args: structured plan · auto-fill affected artifacts from Impact Graph · AI plan critique · scope guard prevents out-of-list edits.
🧪 Lab · experiments & decisions
🧪 Hypothesis
Business hypothesis → AI-designed experiment → real-traffic verdict.
Problem: business decisions on gut feel · same ideas re-tested every 6 months.
Result: auditable adopt / reject verdict with statistical evidence.
Args: PICO canvas · sample size & power pre-calc · pre-registered analysis · uses Inspector + Eval + Studio + Cassette snapshots.
🤔 Allocation Adviser new
Before automating: human / AI / deterministic / hybrid — 7-dimension scoring + verdict.
Problem: "default-to-AI" / "default-to-human" anti-patterns.
Result: documented decision per task · auto-sets risk tier · auto-links Studio + Eval.
Args: 7 dimensions · 5 verdicts · workspace-history-aware · monthly cost & throughput estimates.
📡 Operate · daily ops & observability
🔍 Inspector ↪ Langfuse
Permanent record of every run · the backbone every other feature builds on.
Problem: "we can't reproduce" — engineering bottleneck for every complaint.
Result: any complaint resolved <5 min · final artifacts · SOC2/GDPR-ready · pinnable as eval case in 1 click.
Args: per-LLM-call payload · resolved variables · retry/fallback markers · 3-tier retention · PII redaction · feeds Quality Insights, Gate, Usage, Hypothesis evidence.
📈 Quality Insights
Cross-campaign acceptance / regression trends + Hallucination Detector.
Problem: nobody knows when AI quality is degrading until customers complain weeks later.
Result: alerts in minutes · recovery plan with cases auto-captured.
Args: primary + guardrail metrics · auto-correlate with prompt edits · Hallucination Detector · Context Grounding Monitor (6 signals, 4.2× correlation with hallucinations).
🕸 Impact Graph new
Graph of every artifact · answers "if I change X, what breaks?"
Problem: "small" prompt edit silently breaks 12 MCT tenants · contracts have invisible consumers.
Result: 3-hop blast radius with red / amber / green edges · Change Plan pre-fills from impact.
Args: prompt ↔ DAG ↔ eval ↔ contract ↔ KB ↔ campaigns · killer-query box · saved monitors ("prompts with stale KB").
💬 Customer Feedback new
👍 / 👎 + comments from end-customers · triage · pattern detection.
Problem: only implicit signal — actual recipient voice never reaches the iteration backlog.
Result: 👎 → pattern → "Add to Eval" or "Open Hypothesis" · 1-click reply with evidence link.
Args: NPS / CSAT · back-link to Inspector run · pattern detection across complaints · triage queue.
💳 Usage Hub
Budgets · chargeback · BYOK · Usage Advisor (cheap-model swaps).
Problem: engineering-only cost view · finance can't attribute to clients.
Result: per-team budgets with soft/hard stops · chargeback PDF for finance · Advisor captures $78/mo by classifier Opus→Haiku-style swaps.
Args: tag-based attribution from Inspector · BYOK (we're out of billing path) · Usage Advisor grounded in cassette evidence.
📅 Weekly Review
Auto-prep Monday deck · 7 fixed cards · action items with owners.
Problem: rituals drift · "what did we agree last week?" takes 15 minutes.
Result: Sunday 22:00 auto-prep pulls from Quality · Reviews · Hypothesis · Usage · Feedback · KB. Ritual health tracked too.
Args: 7 action-oriented cards · owners + deadlines set in meeting · rolling 12-week trend answers "are we getting better".
⚙ Settings · workspace configuration
Always-On AI ↪ Portkey
Policy-driven multi-provider fallback + simulator.
Problem: "Sonnet is down → Acme Q2 drops" · blast radius is a page after the fact.
Result: per-tier fallback policy · cassette-based drills · zero production impact when rehearsing.
Args: primary → fallback chain per risk tier · retry / circuit-breaker · graceful-degrade · each fallback logged with cost overhead tag.
🧠 Model Catalog new
Unified catalog across providers · pricing · sunsets · BYOK · side-by-side compare.
Problem: model lifecycle drift · nobody tracks which prompts depend on a sunset-eligible model.
Result: deprecation auto-flags every prompt using a sunset model · migration hypothesis auto-generated.
Args: per-model pricing / p50-p95 / context · side-by-side on real prompts · BYOK 2 of 4 providers · feeds Studio picker, Always-On AI, Usage.
🧬 How features compose
Feedback ↔ Quality a 👎 becomes a pattern, then a regression case captured into Quality Gate — the loop closes from complaint to guardrail.
Cassette ↔ Hypothesis real prod calls recorded → replayed against branches/models without API spend. Adopted hypothesis snapshots the cassette for future replay.
Risk ↔ everything Quality Gate, Reviews, Reliability, Hallucination Detector all consume the resolved tier — classification is load-bearing.
Adviser ↔ Risk verdict auto-sets the tier and creates linked Studio + Eval artifacts. Allocation discipline is not optional.
Impact ↔ Change Plan blast radius pre-fills the plan's affected-artifacts field. Scope guard blocks edits outside the plan.
Weekly Review ↔ all 7 cards aggregate everything — Quality, Reviews, Hypothesis, Usage, Feedback, KB. Ritual health is itself measured.
🏷 Vendor landscape

Buy vs build across 9 categories

22 vendors surveyed · 5 marked as drop-in for our prototype today · the rest stay build.

📚 What this page is for
A walk through the AI-infrastructure market with one question per component: is this still our differentiation, or has it become a commodity we should buy instead of build? Drop-in vendors are flagged with a ↪ vendor pill across the prototype.
Sources: vendor websites · Crunchbase / PitchBook (early 2026) · public customer logos. Tier judgement is opinion, not endorsement.
📌 Summary · what to flag in the prototype today
✓ Already integrated
  • Ragie — RAG / vector store (Chris's branch)
🟡 Drop-in candidates
  • Langfuse → Inspector
  • Braintrust → Quality Gate
  • Portkey → Always-On AI
  • PromptLayer → Prompt Studio
🟢 Build · still our differentiation
  • Output Contract / Intent Validation
  • Hypothesis Lab
  • Usage Hub (cost optimization layer)
  • AI Second Opinion / Coverage Report
  • Impact Graph / Change Plan / Artifact Seal
  • Risk Classification matrix
The asymmetry: 5 commodity layers can be bought (saves engineering quarters); the rest is what makes VelocityEngine recognizable to a customer.
📊 Grafana Cloud · pills across the prototype
Blocks tagged with a Grafana Cloud pill are realizable on our existing subscription with no new infra. Green live pills mark what's already shipped.
✓ Live today
  • Quality Insights — structured-logging dashboard at veprod.grafana.net
🟡 Realizable next
  • Inspector → Cloud Traces (Tempo, OTel LLM)
  • Always-On AI → Cloud SLO · OnCall · Incident
  • Usage Hub → Cloud Metrics + anomaly-detection
  • Hallucination Detector → Cloud Alerting rules
  • Customer Feedback → NPS/CSAT panels
  • Activity timeline → Cloud Logs (Loki)
  • Weekly Review → Cloud Reporting (PDF export)
Why this matters: Grafana Cloud covers the ops infra (traces, metrics, logs, alerts, on-call, SLOs, incidents) — engineering time goes to the campaign-domain layer on top.
🧭 How to read the "buy vs build" call
Maturity tier — Funded Series A+ or OSS standard · multi-tenant · public API + docs · real customer logos. We do not bet on toys.
Replaces real engineering — genuine months-of-work substitute, not a thin wrapper. Building internally would compete with the vendor's full-time team.
Customer can tell the difference — if a customer can distinguish "this is vendor" from "this is you", flag it. If your part is just glue, you are not differentiated.
1 · Retrieval & vector store integrated
Chunking · embedding · multi-tenant retrieval · PDF/DOCX extraction
Vendors at our tier
Ragie in use
Managed RAG with hi-res layout-aware extraction · partition-based isolation · reranker + hybrid search.
Pinecone primitive · DIY chunking
Vector DB only — you own chunking and embedding pipeline. Series B+, enterprise grade.
Weaviate primitive · OSS+cloud
OSS vector store with managed offering · same DIY trade-off as Pinecone.
What we still own
  • Source-of-truth governance · Golden Cases collections · freshness-vs-source tracking · KB drift monitoring
  • Grounding score · 6-signal cross-check
  • Stale-KB consumer alerts · "38 prompts depend on a chunk not re-indexed in 6 weeks"
  • Workflow integration · retrieval as first-class CC-item with type-checked I/O
  • Audit grade · per-run RAG_QUERY events with chunk-level provenance
Verdict: ✅ Buy. Done. Chris has the integration in branch RAG-ccItem.
2 · LLM observability / tracing flagged
Per-call payload · resolved variables · cost · latency · retry markers · session grouping
Vendors
Langfuse recommended
OSS + cloud · ~10k GitHub stars · $4M YC W23 · Samsara, Khan Academy on logos. De-facto OSS standard.
Helicone YC W23 · proxy
Proxy-style observability + caching. 2k+ companies. Lower friction, higher per-call latency overhead.
Arize Phoenix / AX $70M Series B
OTel-native · Uber · ServiceNow. Heavyweight — overkill if you only need per-call inspector.
What we still own · Inspector
  • Business-context tagging · campaign ID · asset type · MCT lineage. Vendor sees calls; we see which campaign for which client
  • Domain events · SEAL_VERIFIED · OUTPUT_CONTRACT_OK · INTENT_VALIDATED · RAG_UPLOAD_FAILED · ARTIFACT_SEAL_BROKEN. Our policy primitives, not observability
  • "Reproduce in <5 min" · Inspector → Studio → Eval Gate → Customer Feedback — that's a product, not a trace viewer
  • Audit-grade retention · 3-tier with PII redaction tied to risk classification
Verdict: 🟡 Strongest "buy" candidate. Frees engineering to focus on the campaign-context layer.
3 · Eval-as-a-Service flagged
Datasets · LLM-judge calibration · regression detection · CI integration
Vendors
Braintrust recommended
$36M Series A (a16z) · Notion · Stripe · Airtable · Zapier. Vocabulary closest to our Quality Gate.
Langfuse Evals bundled
Evals + dataset versioning + judge templates inside the same product as observability. One-vendor story is attractive.
Patronus AI $17M seed
Pre-built evaluators (hallucination, PII, toxicity) · good complement, not a Quality Gate replacement.
What we still own · Quality Gate
  • Campaign-quality rubric · brand voice · promo correctness · regulated-industry tone. Vendors don't know what "good copy for Acme Bank" looks like
  • Coverage Report · 6-dimension matrix telling where the suite is weak, not just which cases pass
  • Auto-capture · from Inspector + Customer Feedback — complaint → permanent regression test
  • Publish-gate enforcement · sits in the Studio publish flow with seal-and-revert mechanics
Verdict: 🟡 Strong "buy". Eval orchestration is commodity now.
4 · Prompt management flagged
Versioning · branches · A/B · labels · rollback · non-engineer editing
Vendors
PromptLayer recommended
Production at ENS · Gorgias · Speak. Established 2022 · mature non-engineer editing UI.
Langfuse Prompts bundled
Versioned prompts integrated with tracing. Single-vendor consolidation play.
Latitude OSS · young
OSS prompt platform with versioning + evals. Smaller ecosystem.
What we still own · Studio
  • Risk tab · Output Contract · Intent Validation · runtime policy declared per CC-item
  • Change-Plan-gated edits · can't open Studio for production prompts without an approved plan in scope
  • Impact Graph reverse-deps · "this edit breaks 12 MCT tenants" — vendor doesn't see your DAG
  • Domain linter · contract violation risk · role-tagging on user data · KB usage hygiene
  • Multi-stakeholder review · AI Second Opinion + diligence checklist + reviewer reliability score
Verdict: 🟡 "Buy" storage / versioning. Keep our authoring UX wrapping it.
5 · LLM gateway / failover flagged
Multi-provider routing · retry · circuit breaker · cost cap · semantic caching
Vendors
Portkey recommended
$3M seed · production at Postman · Springworks · 200+ models · cleanest gateway UX.
OpenRouter ~$100M ARR · routing-only
Massive dev adoption · pure routing without observability/policy depth.
LiteLLM OSS · $6M
OSS proxy + hosted · 100+ providers · virtual keys · budgets. Adobe · Lemonade as users.
What we still own · Always-On AI
  • Per-CC-item policy · "MULTI_LAYER prompts must use claude-opus-4.7, no fallback" · vendor routes, we set policy
  • Tier-aware enforcement · Risk Classification decides what fallback is allowed per tier
  • Compliance guards · Require Exact Model + region restrictions tied to client contracts
  • Reliability event correlation · fallback events surface in Inspector with cost overhead tagged
Verdict: 🟡 "Buy" the plumbing. Keep tier-aware policy as our governance layer.
6 · Guardrails / output validation careful
JSON schema · PII · topic filtering · structured output · hallucination check
Vendors
Guardrails AI $7.5M seed · OSS+cloud
Hub of validators · ~4k stars. Generic validators, not domain rules.
Patronus AI $17M seed
Runtime guardrails · hallucination · retrieval relevance · custom policies · enterprise focus.
NeMo Guardrails NVIDIA · OSS
Colang DSL for dialog flow · self-host.
Why we did NOT flag this
  • JSON schema validation is commodity (Pydantic / Instructor cover it)
  • Output Contract is more than schema — 4-phase lifecycle (write · review · publish · runtime) tied to risk tier and downstream consumer analysis. Vendors don't model that
  • Intent Validation is cheap-classifier-as-gate, lives before the LLM call · vendors validate after
  • Domain rules (no competitor names, no hallucinated SKUs) need workspace-specific policies vendors don't host
Verdict: 🔴 Build. Category is real but immature · "drop-in" framing would mislead the customer.
7 · Feature flags & experimentation flag + our glue
Feature flags · A/B · stats engine · sequential testing · power calculations
Vendors
Statsig $100M Series C · OpenAI customer
Ex-Facebook experimentation team · built-in stats · AI experiment templates.
LaunchDarkly public · enterprise priced
Mature feature flag platform with experimentation. Heavyweight for AI cadence.
Eppo $20M Series A · warehouse-native
DraftKings · Perplexity · Cameo. Stats over your warehouse.
Why we did NOT flag this
  • No vendor is purpose-built for "LLM Hypothesis Lab" — they're general A/B platforms
  • PICO-style framing + AI-designed experiment + pre-registered analysis is our authoring layer · vendor would only run the math
  • Cassette-replay-based experiments (no real prod cost) is unique to our stack
  • Adopted-hypothesis → Change Plan transition is workflow logic, not statistics
Verdict: 🟢 Build now. Consider Statsig later as the stats engine if experiment volume grows.
8 · LLM FinOps / cost observability no winner yet
Per-call cost · per-customer/team rollups · budgets · anomaly detection · forecast
Vendors
Helicone primitive · per-call
Per-call cost attribution + budgets. Useful primitive, not a finance product.
Langfuse Costs primitive · trace-based
Cost per trace/user/session · alerts. Built for engineering, not FinOps.
Vantage / CloudZero FinOps + nascent LLM
Mature FinOps · LLM-specific support just emerging.
Why we did NOT flag this
  • No "Datadog for LLM cost" winner exists yet · everyone builds on Helicone / Langfuse primitives
  • Usage Advisor ("switch this prompt to Haiku → save $78/mo") is a domain-aware optimizer · vendors give numbers, we give recommendations
  • Per-team budgets + chargeback PDF + BYOK live in our workspace model · vendors blind to it
  • Cost-bounded retries tied to risk tier · that's gateway + policy, not a cost dashboard
Verdict: 🟢 Build. Re-evaluate in 12–18 months.
9 · Red-teaming / adversarial testing security, not quality
Jailbreak detection · adversarial prompts · synthetic edge cases
Vendors
Promptfoo ~5k stars · OSS+paid
Red-team plugins · adversarial datasets · CI integration. Shopify · Discord · Anthropic-internal usage.
Lakera $10M seed · enterprise
Runtime + offline red-teaming · "Gandalf" famous · Dropbox · Citi customers.
Patronus AI $17M seed · adversarial sets
Generic red-teaming suites · complement to in-house quality coverage.
Why we did NOT flag this
  • Maturity is real for security red-teaming · we focus on quality coverage
  • AI Second Opinion is workflow-aware (cassette + Coverage Report + contract risk) · vendor red-teams in isolation
  • Coverage Report identifies which dimension is undertested · vendors generate cases without that targeting
  • Promptfoo is interesting as a reference for case-generation patterns · not a drop-in
Verdict: 🟢 Build · quality coverage. Evaluate Lakera if security red-teaming becomes a sales requirement.
Done