Switch workspace
Acme Marketing
Bravo Agency
Charlie Studios
Internal QA Sandbox paused
+ Create workspace
⚙ Workspace settings
SF
🧭
Campaign Explorer — your daily standup screen. See every campaign's health, cost, and recent activity in one place; jump into any feature scoped to one campaign.
why? ▾
📋 What it does
1.Lists every campaign with health (acceptance %), recent runs, cost, alerts.
2.Surfaces a "recommended action" — system correlated alerts with prompt edits.
3.Quick-jump to any feature scoped to the chosen campaign.
🎯 Result
A clear picture of where to spend the next 30 minutes — usually: rollback a regression, approve a teammate's change, optimize a top spender, or send "all-green" to the customer.
🧩 Process gaps closed
·No more "where do I even start?" every morning.
·No more siloed feature dashboards — one campaign-centric view aggregates them.
·Cross-campaign patterns become visible (e.g. all 3 dropping → same prompt).
⚠ Risks mitigated
·Silent regressions discovered late by the customer instead of by you.
·Account-manager surprised at customer review with no data to explain trends.
·Critical campaigns ignored because their dashboard wasn't bookmarked.

Good evening, Sergii 👋

You have 7 active campaigns · 2 need attention · $314 spent on Acme Q2 this week

by status
by client
by team
by tag
flat (no grouping)
Last 24h
Last 7 days
Last 30 days
Custom range…
Workspace summary PDF
All campaigns CSV
Activity log CSV
⭐ Pinned campaigns
Hide ▴
Education Campaign
Needs attention
Acme Q2 Launch
Goals: leads, meetings booked, qualified opportunities
📅 H1-2026 · 47 rejected runs since 14:48 alert
Promo Campaign
Active
Bravo Spring Promo
Goals: brand awareness, demo signups
📅 Q2 · 94% acceptance · $201 spent
B2B Outreach
Running
Golf B2B Outreach
Goals: enterprise SQLs, ABM
📅 Q2 · scheduled run #4787 in progress
Active · 7
Acme Q2 Launch ⚠ 2
Bravo Spring Promo94%
Charlie Holiday91%
Delta Newsletter96%
Echo Black Friday78%
Foxtrot Loyalty93%
Golf B2B Outreachrunning
Drafts · 3
Hotel A/B test
India onboarding
Juliet PR
Archived · 24
▾ Show archived
🚀

Acme Q2 Launch

Needs attention client:Acme · project:Q2

B2B fintech product launch · 12 graphs · DAG with 18 steps · scheduled daily 09:00 UTC

👤 Owner: jess@workspace 🏷 Last published: 2h ago ⏱ Next run: in 5h 24m
Acceptance 7d
87%
▲ 2.1pp
Runs 7d
142
99% ok
Cost 7d
$314
▲ 12%
Avg latency
8.4s
▼ 8%
LLM calls
2.4K
17/run
Open alerts
2
acceptance
Quality & cost — last 30 days
Acceptance Cost Prompt change Alert
100% 80% 70% Mar 21 Apr 1 · prompt edit Apr 11 · prompt edit Apr 18 · rollback Today
Composition
Prompts · 6
promo_copywriter tone_classifier json_extractor summarizer +2
KB used · 2
📚 Brand Guides 📚 Product SKUs
Models
claude-3-5-sonnet76%
claude-3-5-haiku18%
gpt-4o (fallback)6%
Resilience policy
Default · fallback to OpenAI
Health summary
Quality⚠ 2 alerts
ReliabilityAll providers healthy
Usage▲ 12% MoM
Eval coverage5/6 prompts
Open reviews2 pending
⚠ Quick actions

Quality dropped 17pp on tone_classifier after a prompt edit at 14:10. Eval suite was not run.

Investigate in Quality Insights →
🔍
Inspector — every run is permanently recorded with the actual prompt sent and the actual answer. When a customer complains, you find the cause in <5 minutes instead of "we can't reproduce".
why? ▾
📋 What it does
1.Persists every step of every run — inputs, outputs, events, errors.
2.For each LLM call: resolved prompt, raw response, model, cost, latency, retries.
3.Search by tag, filter by status, share permalinks.
4."Add to Eval" turns prod case into regression test.
🎯 Result
A finished campaign artifact (email, landing, hero image) ready to ship or send for human review. Plus a full audit trail of every decision the AI made.
🧩 Process gaps closed
·No more "we can't reproduce" responses to support tickets.
·QA team can audit AI behaviour without engineering help.
·Customer support can answer "why did this campaign give X?" themselves.
·Compliance officers get an audit trail by default.
⚠ Risks mitigated
·Loss of customer trust after unexplained bad output.
·Failed compliance audit (GDPR / SOC2 require trace).
·Engineering bottleneck slowing down every quality investigation.
·Repeat regressions because nobody captures bad cases as tests.

🔍 Inspector Langfuse Grafana Cloud Traces

142 runs in the last 7 days · 3 failed · $314 cost · scope: Acme Q2 Launch

All
Success
Partial
Failed
Cancelled
Running
Last 24h
Last 7 days
Last 30 days
This quarter
Custom range…
Filter by tag
☐ client:Acme
☐ client:Bravo
☐ client:Charlie
☐ project:Q2
☐ internal
Export CSV (current view)
Export JSON (with payloads)
Generate PDF report
Showing 12 of 142 in scope
#4823 · 14:32
generate_promo_assets
$1.20
12.4s
#4828 · 14:18
generate_promo_assets
$0.40
47.1s
#4819 · 14:02
generate_promo_assets
$0.05
3.2s
#4811 · 13:30
generate_promo_assets
$1.18
11.1s
#4807 · 13:01
generate_promo_assets
$1.22
12.8s
#4803 · 12:30
generate_promo_assets
$1.15
10.9s
#4799 · 12:01
generate_promo_assets
$1.21
11.7s
#4795 · 11:30
generate_promo_assets
$0.62
22.4s
#4793 · 11:14
⊘ intent rejected · legal
$0.001
0.4s
#4712 · 09:48
🚫 contract violation · MCT
$0.04
2.1s
#4708 · 09:32
🔓 seal broken · agent edit
$0.00
paused
#4694 · 08:51
⚠ rag_upload failed · retry
$0.18
8.4s
#4791 · 11:01
generate_promo_assets
$1.19
11.4s
#4787 · 10:30
generate_promo_assets
$1.16
11.0s
#4823 success generate_promo_assets 14:32 · 12.4s · 24 LLM calls · $1.20 · 42.1K tokens
Add tag
client:Acme
client:Bravo
project:Q2
priority:high
+ New tag…
Same inputs, current prompts
Same inputs, prompt rev 141
Modify inputs first…
🔀 Open in Workflow
0s3s6s9s12.4s
✨ Quick actions
This run succeeded but generated output flagged by Hallucination Detector (claim "Acme holds $2.3B AUM" not in KB). Rebecca got a 👎 from Acme CFO referencing this output. Capture as REGRESSION case + roll back upstream prompt to prevent next-run repeat.
↗ Open prompt in Studio 💬 See Acme feedback
Steps · 18
fetch_briefhttp120ms
kb_lookup_tonekb_lookupRagie1.2s · 3 chunks
extract_intentjson_prompt↻×11.4s · $0.04
apply_toneprompt⇄ openai1.8s · $0.06
generate_promo_textprompt4.1s · $0.42
post_processfunction40ms
generate_dalle_herodalle3.2s · $0.04
moderate_imageguard200ms
compose_emailprompt1.1s · $0.18
compose_landing_htmlprompt1.6s · $0.22
reviewasset2.3s · human
publishassetskipped
Step: generate_promo_text click any step on the left to switch
Inputs
Output
Events · 10
LLM calls · 1
{{brief}}
"Q2 launch for Acme Bank's new B2B fintech product targeting mid-market CFOs in regulated industries."
{{kb_results}}
3 chunks · 1,247 tokens · from GC Brand Guides (rev 7) view chunks
{{tone}}
"professional, warm"
{{brand}}
"Acme Bank"
JSON output (validated against promo.schema.json ✓)
{
  "headline": "Q2 brings precision to your treasury.",
  "bullets": [
    "Real-time cash visibility for CFO teams.",
    "Built for regulated industries — SOC2, HIPAA-ready.",
    "Onboard in days, not quarters."
  ],
  "cta": "Book a 15-min demo"
}
→ Used by next steps: compose_email, compose_landing_html, generate_dalle_hero
14:32:01.118SEAL_VERIFIEDprompt rev 142 · sealed by jess @ CP-101 · hash f3d2…
14:32:01.123STEP_STARTED
14:32:01.140VARIABLES_RESOLVED · 4 vars
14:32:01.148INTENT_VALIDATEDcopywriting · confidence 0.94 · allowed ✓
14:32:01.215RAG_QUERYretrieve_brand_voice · "tone for Acme treasury" · topK=8
14:32:01.487RAG_QUERY_OK3 chunks · top score 0.91 · brand_guides_acme rev 7 · 272ms
14:32:01.155BUDGET_CHECK_PASSED · est $0.018
14:32:01.890LLM_CALL_OKclaude-3-5-sonnet · 824ms · $0.018
14:32:01.905OUTPUT_CONTRACT_OKJSON · promo.schema.json · headline:string ✓ bullets:array ✓ cta:string ✓
14:32:05.214STEP_COMPLETED
8f3a2b…OK
claude-3-5-sonnet · attempt 1/3 · $0.018 · 824ms
312 tokens out · valid JSON ✓
→ Full payload visible in the right rail
Final result · #4823
3 assets generated · ready for review
📧 Email
Q2 brings precision to your treasury.
Real-time cash visibility for CFO teams. Built for regulated industries — SOC2, HIPAA-ready. Onboard in days, not quarters.
subject_v1.html · 4.2 KBOpen ↗
🌐 Landing
Acme Treasury · Q2 launch
Hero + 3 benefit cards + CTA "Book a 15-min demo". Tailwind, mobile-ready.
landing_v1.html · 18.2 KBOpen ↗
🖼 Hero image
Acme Treasury
DALL·E · 1024×1024 · ✓ moderatedFull ↗
Total cost $1.20 · 12.4s · published to staging bucket pending review
LLM Call detail
8f3a2b…
Meta
Request
Response
Trace
Model
claude-3-5-sonnet
Cost
$0.018
Latency
824 ms
Tokens (in/out)
1,243 / 312
Temperature
0.7
Max tokens
1,024
Top P
1.0
Attempt
1/3
Cache
Fallback
Prompt rev
142 diff
Provider req
req_01J9F4B2…
Resolved prompt
You are a copywriter for Acme Bank.
Tone: professional, warm.

Brand context (KB):
[3 chunks · 1,247 t]

Brief:
Q2 launch for Acme Bank's new B2B fintech product…

Output a punchy headline + 3 supporting bullets in JSON.
Response · valid JSON ✓
{
  "headline": "Q2 brings precision to your treasury.",
  "bullets": [
    "Real-time cash visibility for CFO teams.",
    "Built for regulated industries — SOC2, HIPAA-ready.",
    "Onboard in days, not quarters."
  ],
  "cta": "Book a 15-min demo"
}
14:32:01.123request sent → Anthropic api.eu
14:32:01.211connection established (88ms)
14:32:01.890first byte (767ms)
14:32:01.947complete (824ms)
14:32:01.952JSON validated against schema
↗ Open in Studio
Pin to suite
promo_copy_quality (47 cases)
json_extractor_check (120)
tone_classifier (30)
+ New eval suite…
🔀
Workflow Editor — visual graph of CC-items (prompts, KB lookups, JSON parsers, HTTP calls, image gen, human approval). The actual "no-code pipeline builder" customers use to assemble a campaign.
why? ▾
📋 What it does
1.Drag CC-items from the palette onto the canvas
2.Connect outputs → inputs (typed contracts; mismatches highlighted)
3.Per-node inspector: Prompt Studio for prompts, KB picker for kb_lookup, etc.
4.Live cost & latency estimate per run as graph grows
5.Validation: cycles, unconnected vars, missing risk tier on MULTI_LAYER
🎯 Result
A shippable campaign DAG — versioned, testable, with cost+latency known before first run. The thing that actually generates the customer's promo email + landing + image.
🧩 Process gaps closed
·Non-engineers assembling pipelines without writing code
·"Where exactly is this prompt used?" answered visually
·Graph-level review (not just per-prompt) before publish
⚠ Risks mitigated
·Bad graph topology causing infinite retry loops
·Variable type mismatches caught at design time, not at runtime
·Unguarded MULTI_LAYER outputs → linter blocks save

🔀 Workflow Editor

Acme Q2 Launch · generate_promo_assets · 18 nodes · cost preview $1.20/run · 12.4s avg

🔒 14 sealed 🔓 3 broken (in CP-104) — 1 unsealed (TRUST tier) 🕸 see blast radius for this DAG 📋 active Change Plan

⎇ main (production)
⚡ experiment-tone-v2 (Jess · draft)
⎇ holiday-variant (Sergii · in review)
Palette · drag to canvas
LLM nodes
⋮⋮promptPlain prompt
⋮⋮json_promptJSON-schema prompt
⋮⋮image_promptImage gen
⋮⋮dalleDALL·E (MULTI_LAYER)
⋮⋮session_promptMulti-turn chat
Context nodes
⋮⋮kb_lookupGolden Cases retrievalRagie
⋮⋮retrievalVector DB queryRagie
⋮⋮web_scraperURL → text
⋮⋮file_contentLoad fileRagie ext.
⋮⋮httpREST call
Logic / data
⋮⋮functionJS / Python
⋮⋮jsonJSONPath / jq
⋮⋮regexPattern extract
⋮⋮foreachLoop over array
Quality gates
⋮⋮guardSchema validator
⋮⋮moderateContent moderation
⋮⋮human_approvalPause for human
⋮⋮assetSave artifact
Canvas
18 nodes · 24 edges·zoom 100%
START httpfetch_brief kb_lookuptone_chunks json_promptextract_intentselected promptapply_tone promptgenerate_promo$0.42 · 4.1s functionpost_process promptcompose_email promptcompose_landing dallegenerate_heroMULTI_LAYER promptcompose_subject guardmoderate_email moderatemoderate_image guardmoderate_landing human_approvalreview & approveRebecca assetpublish to micrоsitestaging-bucket
⚠ Linter findings · 2 (graph-level)
  • generate_hero (MULTI_LAYER) — moderate_image is connected, but no human_approval edge before publish · add edge
  • extract_intent outputs {intent, audience} but downstream apply_tone only consumes tone — unused output · view types
⚠ Cannot publish yet
Graph has 1 CRITICAL linter finding (MULTI_LAYER node generate_hero can publish without human approval). Auto-fix is available: insert human_approval node between moderate_image and publish.
Selected node
json_promptextract_intent
Risk tier: HUMAN_REVIEW · Cost ~$0.04/run · 1.4s avg
Inputs · 2
{{brief}}from fetch_brief
{{kb_results}}from kb_lookup
Outputs · 2
intent : enum→ apply_tone
audience : stringunused ⚠
Run estimate · whole graph
Per-run cost$1.20
Per-run latency~12.4s
LLM calls24
Tokens (in/out)~42K / 8K
For 100 runs/day: ~$3,600/mo
🧠
Model Catalog — single place where you see every LLM model available to your workspace, what it costs, who uses it, and BYOK keys. Default model selection per use-case (draft / review / final).
why? ▾
📋 What it does
1.Lists all available models from connected providers
2.Per-model: cost / latency / context window / capabilities
3.Workspace defaults: draft → Haiku, review → Sonnet, final → Opus
4.BYOK key management per provider
5.Deprecate / disable specific models (sunset planning)
🎯 Result
One place to govern what AI you allow in your workspace, who pays for it (BYOK vs platform), and what defaults Studio offers when authoring.
🧩 Process gaps closed
·"Which models can I use?" — clear answer
·Compliance: only approved models in production
·Sunsetting deprecated models with migration plan
⚠ Risks mitigated
·Surprise model deprecation by provider mid-quarter
·Procurement blocks deal because no BYOK
·"Wait, when did we start using gpt-5-preview in prod?"

🧠 Model Catalog

9 models from 4 providers · 2 BYOK keys configured · 1 deprecation upcoming

Providers · 4
Anthropic3 models
OpenAI3 models
Google AI2 models · degraded
Bedrock1 model
Workspace defaults
draftclaude-3-5-haiku
reviewclaude-3-5-sonnet
finalclaude-3-opus
judge (eval)gpt-4o-mini
summarizerhaiku-3.5
imagedall-e-3
Filters
⚠ Quick actions
claude-3-opus deprecates Jul 31, 2026 (102 days). 2 prompts (340 calls/wk) will break. Migration target ready: claude-opus-4.7. Run shadow A/B now via Hypothesis Lab — auto-creates Cassette baseline + paired eval. Verdict in ~2 weeks.
🧪 Run shadow A/B on opus-4.7
Anthropic models · 3
used: 89% of LLM calls this week
ModelStatusCost (in/out)p95 latencyContextUsed byCalls 7d
claude-3-5-sonnet-20241022 default · review● active$3 / $15920ms200K14 prompts8.2K
claude-3-5-haiku-20241022 default · draft● active$0.80 / $4410ms200K5 prompts2.1K
claude-3-opus-20240229 final / sunset Q3● deprecating Jul 31$15 / $752.1s200K2 prompts340
OpenAI models · 3used: 8% (mostly fallback)
ModelStatusCost (in/out)p95 latencyContextUsed byCalls 7d
gpt-4o-2024-11-20 fallback● active$2.50 / $101.4s128Kvia Always-On421
gpt-4o-mini judge● active$0.15 / $0.60680ms128Keval judges214
dall-e-3● active$0.04 / image3.2s1 prompt26
Google AI · 2Degraded · CB half-open
ModelStatusCost (in/out)p95 latencyContextUsed byCalls 7d
gemini-1.5-pro⚠ degraded$1.25 / $54.2s2M1 prompt12
gemini-1.5-flash● active$0.075 / $0.30540ms1M0
🔑 BYOK keys · 2 of 4 providers
Anthropic your contractconnected
OpenAI your contractconnected
Google AI platform-billedadd key
Bedrock platform-billedadd key
⚠ Sunset planning
claude-3-opus-20240229 deprecates Jul 31, 2026.
Migration target: claude-opus-4.7 (preview). Affects 2 prompts (340 calls/wk).
Spotlight
claude-3-5-sonnet
2024-10-22 · workspace default for review
Context200K tokens
Vision
Tool use
JSON mode
Caching✓ 90% off cached
Streaming
EU residency✓ via api.eu
Used by · 14 prompts
Compliance
Approved for production
SOC 2 Type II✓ provider
GDPR DPA signed
HIPAA BAAon request
💬
Customer Feedback Inbox — explicit thumbs-up/down + comments from end-users (the people who actually consume your generated copy/images). Triage queue; convert into eval cases or hypothesis seeds.
why? ▾
📋 What it does
1.Aggregates 👍/👎 + comments from chat-embed widgets & landing pages
2.Triage queue: new / acknowledged / resolved / converted-to-eval
3.NPS + CSAT trends per campaign
4.One click: "Add as eval case" / "Open hypothesis"
5.Notification routes (Slack thread per negative)
🎯 Result
Direct line from end-customer voice to your iteration backlog. Not just implicit "they didn't accept" — actual "tone is too pushy" or "you said $20K but our deal is $5K".
🧩 Process gaps closed
·End-customer never had a say in AI quality before
·Sales team forwards complaints — now they go to triage queue
·Quality stays anchored to actual recipient experience
⚠ Risks mitigated
·Quality drift visible to customer before to you
·Pattern of complaints (e.g. tone) not aggregated
·NPS drop discovered in QBR instead of week-of

💬 Customer Feedback Grafana Cloud Metrics

312 signals this week · NPS 47 (▲ 4) · 14 new in queue · 3 converted to eval cases

New
Acknowledged
Resolved
Converted to eval
All
24h
Last 7d
Last 30d
Inbox · 14 new
👎Acme CFO14:42
"This sounds way too informal — we audit Big Four clients."
acme-q2tone
👎Bravo lead13:18
"Headline was great but CTA didn't match offer."
bravo-springcta-mismatch
👍Charlie ops12:50
"Loved the email subject — opened immediately."
charlie-holiday
👎Acme procurement11:33
"You said $2.3B AUM — that's not us, we're private & smaller."
factualacme-q2
👎Echo manager10:42
"Image looked stock — generic, not our brand."
echo-bfbrand-image
👍Delta editor09:51
"Nailed the seasonal angle, perfect."
delta-newsletter
+ 8 more in queue
Selected · Acme CFO · 14:42
"This sounds way too informal — we audit Big Four clients. Tone is fine for marketing but I can't put my name on this email going to enterprise CFOs."
Source
📧 Email widget on landing page
Recipient
CFO · acme-bank.com
🔍 Pattern detected
This is the 3rd "tone too informal" complaint this week, all about tone_classifier rev 142. Suggested action: view Quality Insights alert · or open Hypothesis "tone calibration for B2B finance".
NPS (4 weeks)
47
▲ 4
CSAT
4.3 / 5
▲ 0.2
👍 / 👎 ratio
82%
▼ 3pp · tone issue
Top complaint themes · last 7d
Tone too informal
8
CTA mismatch
4
Factual error
3
Brand-off image
2
Length / wordiness
1
Per-campaign sentiment
Campaign👍👎RatioTrend
Delta Newsletter42295%▇▇▇▇▇▇
Bravo Spring38588%▆▇▆▇▇▇
Acme Q2521281% ▼▇▇▇▆▅▄
Charlie Holiday28390%▇▇▇▇▇▇
Echo Black Friday21872% ▼▆▅▄▄▃▃
Triage actions
Recently converted · 3 this week
"too pushy"tone_classifier eval
"wrong AUM number"factual hallucination eval
"image looks stock"DALL·E brand-safety hypothesis
Embed widget

Drop this snippet on customer-facing surface (email, landing, chat) to collect feedback:

<script src="https://cdn.velocityengine.io/feedback.js"
   data-workspace="acme-marketing"
   data-campaign="acme-q2"></script>
📈
Quality Insights — answers "is the AI getting better or worse?" with real numbers. Detects drift, finds the cause, lets you roll back in one click — before the customer notices.
why? ▾
📋 What it does
1.Tracks acceptance ratio (humans accepting AI output) per campaign / prompt / model.
2.Detects drift vs baseline, fires alerts.
3.Correlates drops with prompt/model edits — shows likely cause.
4.One-click rollback to last working version.
🎯 Result
A trustworthy quality dashboard for OWNER + a recovery plan when something breaks: 5-step recovery, captured cases, audit log, customer-shareable PDF.
🧩 Process gaps closed
·No more "did this change make things worse?" guessing.
·No more weeks-long delay between quality drop and discovery.
·Sales reviews with customer have data, not anecdotes.
⚠ Risks mitigated
·Silent quality decay → customer churn.
·Reputational damage from undetected output regressions.
·Lost insight into which prompt/model combos work — blind future edits.

📈 Quality Insights Grafana Cloud · live

Acceptance 87.3% · 2 open alerts · 47 rejected runs since 14:48

Last 24h
Last 7 days
Last 30 days
This quarter
YTD
Aggregate by
Workspace
Campaign
Prompt
Model
Team
Cost center
Format
PDF · branded
PDF · executive 1-pager
CSV · raw aggregates
Schedule
Send weekly to client@…
Send monthly to OWNER
Open alerts · 2
Acceptance ↓ 17pp
tone_classifier · 38m ago
Acme Q2 · Bravo Spring
Refusal rate ↑ 5×
json_extractor · 2h ago
Echo Black Friday
Recent (resolved) · 4
Schema fail spike · 6h ago · auto-resolved
Acceptance ↓ summarizer · 1d ago
Latency ↑ p95 · 2d ago
Acceptance ratio
87.3%
▲ 2.1pp · 4,210 decisions
Schema violations
1.2%
▲ 0.4pp
Refusals
0.4%
▲ 5×
Eval coverage
5/6
prompts with suites
Acceptance ratio · last 30 days
Acceptance Prompt change Alert Rollback
100% 80% 70% Mar 21 Apr 1 Apr 11 Apr 18 Today
By model
claude-3-5-sonnet90%
claude-3-5-haiku86%
gpt-4o (fallback)82%
claude-3-opus94%
Acceptance heatmap (24h × 7d)
SunSat
🧠 Hallucination Detector · last 7d Grafana Cloud alertingcross-references model output against GC chunks + input data + JSON-citable claims
Outputs scanned
2,140
100% of MULTI_LAYER + REVIEW
Suspect claims
23
1.07% rate
Confirmed hallucinations
4
0.19% · all blocked
Avg grounding score
0.91
target ≥0.8
Recent confirmed hallucinations · 4
TimeRunClaimSource checkAction
Apr 18 16:22#4712"Acme Bank holds $2.3B AUM"not in GC · not in briefblocked
Apr 17 11:09#4604"audited by Big Four"contradicts brief: "Acme is private"blocked
Apr 16 09:44#4521"40% YoY growth"brief says "20% YoY"blocked
Apr 15 14:31#4480"NYSE-listed"not stated · public web check failsblocked
Detection method: AI claim-extractor pulls factual claims from output → matches against GC embeddings Ragie, input variables, and (optionally) live web search.
Action by tier: TRUST = warn-only · REVIEW = warn + log · CO-CREATION = warn user · MULTI-LAYER = block + reroute to human
🧭 Context Grounding Monitor · last 7dtracks freshness, retrieval quality, drift & budget pressure of context that lands in the model window
Avg grounding score
0.87
target ≥0.8
Low-grounding runs
112
5.2% · score <0.6
Stale-KB runs
38
KB rev >30d behind source
Budget-truncated
19
prompt >90% of model window
Signal breakdown · how grounding score is computed
KB freshness (rev age vs source)0.78
Retrieval relevance (top-k score)0.91
KB drift (chunk-set churn)0.84
Memory recall (long-context loss)0.72
Token budget headroom0.88
Variable resolution coverage0.95
Top stale-context offenders · 4
KB / SourceLast syncSource updatedAffected campaignsAction
brand_guides rev 7Mar 4Apr 12 · 6 sections changedAcme Q2, Bravo SpringRe-sync
compliance_rules_eu rev 12Feb 28Apr 10 · GDPR addendumCharlie Holiday EURe-sync
tone_examples rev 4Apr 1Apr 18 · Jess added 12Echo Black FridayRe-sync
product_catalog rev 22Apr 17Apr 19 · 3 SKUs addedFoxtrot LoyaltyRe-sync
Action by tier: TRUST = log only · REVIEW = warn in run · CO-CREATION = require fresh-KB confirm · MULTI-LAYER = block run if score < threshold (default 0.6).
How the score helps: low grounding correlates 4.2× with hallucinations and 2.7× with 👎 feedback — fixing freshness is upstream of both detectors.
💬 ChatOverride · accept/reject signal · last 7dhow often humans accept AI suggestions in chat-style flows
AI suggestions
1,847
across CO-CREATION nodes
Applied
1,512
81.9% accept rate
Canceled
201
10.9% rejected
Edited & applied
134
7.2% partial
Accept rate trend (30d)
90% 80% 70% Mar 21 Apr 1 Apr 11 Apr 18
Per CC-item leader:
chat_assistant_acme · 87% accept
Per CC-item laggard:
chat_holiday_promo · 64% (needs attention)
Most-edited suggestion:
compose_subject_v2 · 28% edit-then-apply
Selected alert
Acceptance ↓ 17pp · tone_classifier
Observed 71% · Baseline 88% · Threshold −15%
Detected 14:48 · Affected: 3 campaigns · 47 rejected runs
Likely cause

Prompt tone_classifier was edited 38m before the drop.

Editor: jess@workspace · Branch: main (no Eval Suite ran)
Open Studio
Distribution shift

Inputs to this prompt drifted: queries about "compliance" rose 3.2× this week vs eval set.

→ Add cases to eval suite
Recent rejected runs · 47
14:47 · #4831 · LLM_JUDGE failed
14:42 · #4828 · schema fail
14:38 · #4823 · user canceled
14:31 · #4819 · LLM_JUDGE failed
14:24 · #4819 · LLM_JUDGE failed
→ Open all in Inspector
→ See related Customer Feedback (3 patterns)
Always-On AI — when Anthropic / OpenAI / Google have a bad day, your campaigns keep running by automatically falling back to a healthy provider. No 3am pages, no missed customer deadlines.
why? ▾
📋 What it does
1.Monitors provider health (latency, error rate, circuit-breaker state).
2.Defines policies: primary → fallback chain (drag to reorder), retry rules.
3.Runs failure simulations dry — see what would happen before it happens.
4.Routes alerts to email / Slack / PagerDuty.
🎯 Result
Campaigns that never go dark during a provider outage. Every fallback is logged in Inspector with a marker so you can review what happened.
🧩 Process gaps closed
·No more single-provider dependence on Anthropic or OpenAI.
·No manual scramble to switch providers during incidents.
·Engineering on-call doesn't get paged for routine 5xx bursts.
⚠ Risks mitigated
·Customer SLA breach during a 30-minute Anthropic outage.
·Retry-storm against a degrading provider amplifying the outage.
·Bill shock from runaway retries (cost guard caps it).
·Compliance violation from silent provider switch (Require-exact-model toggle prevents it).

⚡ Always-On AI Portkey Grafana Cloud SLO · OnCall · Incident

4 policies active · Anthropic + OpenAI healthy · Google AI degraded · 47 fallbacks today (+$8.40)

Policies · 4
Default
used by ~all CC-items
Strict — no fallback
4 CC-items · compliance
Cheap with brownout
12 CC-items
Multi-provider HA
1 CC-item · critical
Alert routes
📧 ops@acme.com
💬 #ai-incidents (Slack)
📟 PagerDuty (critical only)
Anthropic
Healthy
CB closed · 0% failure (5min) · p95 920ms
OpenAI
Healthy
CB closed · 1.2% failure (5min) · p95 1.4s
Google AI
Degraded
CB half-open · 23% failure (5min) · p95 4.2s
⚠ Quick actions
Google AI degraded (23% failure · CB half-open · p95 4.2s). Currently it's in Multi-provider HA policy as Tier 2. Either drop it from chain until recovery, or add Bedrock as backup tier so HA stays redundant. Live calls already auto-routing to other tiers.
Editing: Default policy
Primary
Fallback chain · drag to reorder
⋮⋮ TIER 1 OpenAI · gpt-4o cost ×1.2
⋮⋮ TIER 2 Google · gemini-1.5-pro cost ×0.9
⋮⋮ TIER 3 Cached response TTL 1h · triggers ALL
Provider fallback
Anthropic / claude-3-5-haiku
OpenAI / gpt-4o-mini
Bedrock / claude-3-sonnet
Vertex / gemini-1.5-flash
Cached response (any TTL)
Static fallback message
Route to human queue
Retry
Max attempts
Initial backoff
Multiplier
Jitter
Circuit breaker
Failure threshold
Open duration
Window
Timeout
Compliance & cost guards
Live preview
Anthropic / sonnet
↓ on 5xx / timeout / rate_limit · 3 retries · backoff 500ms ×2
OpenAI / gpt-4o (×1.2)
↓ on circuit_open
Cached response (1h)
↓ on cache_miss
✗ User-facing failure
Max attempts before failure: 9 (3 primary × 3 tiers). Worst-case latency budget: ~95s.
Reliability events · last 24h
5 fallbacks 23 retries 1 outage
TimeEventProviderDetailRun
14:42⇄ fallbackAnthropic → OpenAI5xx (Sonnet 503) · 3 retries exhausted#4828
13:18↻ retryAnthropicrate_limit 429 · backoff 500ms#4819
12:42⇄ fallbackAnthropic → OpenAI5xx burst · 3 retries exhausted (today's primary incident)#4811
11:00⊘ CB openGoogle AIfailure rate 67% > 50% threshold · open 60s
10:15↻ retryOpenAIschema_violation · re-prompt#4799
Simulator

Dry-run a failure to see what your policy would do.

Active alerts
Google AI degraded
CB half-open · 23% failure
Notified: ops@acme.com
Provider quotas (today)
Anthropic847 / 5,000 RPM
OpenAI2.1K / 10K RPM
Google AI120 / 1K RPM
✏️
Prompt Studio — a real IDE for the people who write prompts. Branches like git, side-by-side A/B playground, live cost estimates, an inline linter that catches injection risks and informal output specs before they ship.
why? ▾
📋 What it does
1.Edit prompts on a draft branch, never directly in production.
2.Drag variable chips into the editor; see resolved values.
3.Compare A/B output side-by-side in playground.
4.Inline linter flags injection risks, missing schemas, unused vars.
5."Publish" triggers Quality Gate — eval suite must pass.
🎯 Result
A reviewed, version-controlled prompt that has been tested before going live and that someone signed off on. With cost & latency estimates per model already known.
🧩 Process gaps closed
·No more "I'll just edit prod and see what happens".
·No more lost edits — versions tracked, rollback in 1 click.
·No more reading raw config to understand a prompt's variables.
·Non-engineers can author prompts without breaking things.
⚠ Risks mitigated
·Prompt injection from user data (linter wraps it).
·Production breaking on save with no review.
·"Who changed this?" — full author + reason history.
·Picking the wrong model — cost preview shows true monthly $.

✏️ Prompt Studio PromptLayer

Editing promo_copywriter @ experiment-tone-v2 · 2 linter warnings · cost preview $0.018/call on Sonnet

🔓 seal broken on draft edit gated by approved plan CP-101 · intent allowlist (2d ago, by jess) · in-scope artifacts: promo_copywriter, compose_email 🕸 see 11 downstream
promo_copywriter @
Switch branch
⎇ main production
⚡ experiment-tone-v2 draft
⎇ holiday-variant in review
+ New branch from main
+ New branch from current
DRAFT · unsaved changes
▶ Run all eval cases (47)
▶ Run regressed-only (3)
▶ Run on single input…
▶ Shadow A/B on prod (24h)
Prompts · 12
📄 promo_copywriter draft
📄 tone_classifier
📄 json_extractor
📄 summarizer
📄 brief_extractor
📄 compose_email
📄 compose_landing_html
+ 5 more
Branches · 3
⎇ main production · rev 142
⚡ experiment-tone-v2 +12 −8
⎇ holiday-variant in review
Version history
rev 142 · current
you · just now
rev 141 · main
you · 2h ago
rev 140
maria · 1d ago
rev 139
ivan · 3d ago
+ 138 more in Envers
✨ Before you publish
You have unsaved changes on branch experiment-tone-v2, 2 linter warnings open, and Quality Gate hasn't run yet on this revision. Apply linter fixes + run eval against promo_copy_quality (47 cases · ~$0.85) before requesting review.
Editor
1,243 chars · 312 tokens · Last save: just now (auto)
1You are a copywriter for {{brand}}.
2Tone: {{tone}}.
3
4Use insights from the Golden Cases library:
5{{kb_results}}
6
7Brief:
8{{brief}}
9
10Output a punchy headline + 3 supporting bullets in JSON.
11// Linter: line 10 — output format spec is informal, consider JSON Schema
12↪ Drop variable chips from right rail to insert
Playground
Today: 14 runs · $0.32
Inputs
{{brand}}
{{tone}}
{{brief}}
Or load from eval case ▾
⚡ experiment-tone-v2 · $0.020 · 880ms
{
  "headline": "Yo, Acme Bank's got your treasury covered.",
  "bullets": [
    "See cash in real-time, anytime.",
    "Compliance-ready out the box.",
    "Up and running in days."
  ]
}
⚠ LLM_JUDGE: tone score 3/10 (rubric: professional)
⎇ main · $0.018 · 824ms
{
  "headline": "Q2 brings precision to your treasury.",
  "bullets": [
    "Real-time cash visibility for CFO teams.",
    "Built for regulated industries — SOC2, HIPAA-ready.",
    "Onboard in days, not quarters."
  ]
}
✓ LLM_JUDGE: tone score 9/10
Vars
Linter 3
Cost
Used by
Memory
Risk
Variables · drag to insert
⋮⋮{{brand}}
Source: brief · "Acme Bank"
⋮⋮{{tone}}
Source: env_var · "professional, warm"
⋮⋮{{kb_results}}
kb_lookup · Brand Guides · ~1,247 t
⋮⋮{{brief}}
intake form · ~1.2 KB
⋮⋮{{client_industry}}
env_var · "fintech"
⋮⋮{{cta}}
env_var · "Book a 15-min demo"
Tip: drag a chip onto the editor area to insert a reference.
Eval coverage
promo_copy_quality47 cases · 92%
→ Open suite
Linter findings · 3
🚫 Line 8 — output contract violation risk
Prompt instructs "return a list of headlines" but downstream compose_email calls headline.toLowerCase() (expects string, not array). Will throw TypeError at runtime in <eval> node.
📐 Edit contract
⚠ Line 10 — output format informal
Consider JSON Schema for structured output reliability.
⚠ No role-tagging on user data
Wrap {{brief}} in <user_input> to reduce prompt-injection risk.
ℹ Line 4 — could use Golden Cases v2 reranker Ragie
Brand Guides GC has reranker enabled but query doesn't pass topK.
Cost preview (per call)
Input tokens~1,243
Output tokens (est)~312
claude-3-5-sonnet$0.018
claude-3-5-haiku$0.0025
claude-3-opus$0.092
gpt-4o$0.022
gpt-4o-mini$0.0014
For 1,000 runs/day:
Sonnet $540/mo · Haiku $75/mo · Opus $2,760/mo
→ Full Model Catalog
💡 Usage Advisor
This prompt is a good Haiku candidate based on eval suite scores. Switch could save $465/mo.
Used by · 4 CC-items
Impact analysis
Publishing this change will affect 4 campaigns on next scheduled run. Eval-blocked publishing is enabled — regressions will be caught before going live.
Chat memory · session_prompt mode
Multi-turn prompts pass conversation history to the LLM. Configure how much history is kept and how it is summarized.
Strategy
Window
Last N turns
Max tokens
Summarizer (auto-compression)
Live preview · current session
[system] You are a copywriter for Acme Bank...
[summary, 8 prior turns compressed] User wants Q2 promo for B2B treasury audience, prefers professional+warm tone, wants headline+3 bullets+CTA.
[turn -3] user: "make the headline punchier"
[turn -2] assistant: "Q2 brings precision..."
[turn -1] user: "include compliance angle"
[turn 0 · current] assistant: ...
Memory health
Avg turns per session8.3
Avg context size2,140 tokens
Compression triggered47% of sessions
Avg cost per session$0.034
Risk tier of this prompt
HUMAN_REVIEW
Reviewer must approve before publish · 90d audit · warn-only hallucination detector
Reclassify
Effective controls
Eval suite required≥10 cases · ✓ 47
Pre-publish gateblock on regression · ✓
Reviewer required1 reviewer · ✓ Sergii
Hallucination detectorwarn-only
PII redactionpassive · ✓
Audit retention90 days · ✓
→ View full Risk Classification matrix
Who should do this task?
📌 Allocation Adviser verdict for promo_copywriter:
AI_LED_HUMAN_REVIEW
~3 min review/output · $0.42/output · saves 36h/qtr vs hand-craft
🤔 Open Allocation Adviser →
🛡 Intent validation
Cheap classifier (Haiku · ~$0.0008/call) checks every input before main LLM. Out-of-scope = templated refusal.
Allowed intents · 3
brand_tone_query
brief_clarification
copy_revision
Confidence threshold
0.70
Refusal template
Rejected this week47
Top reasonout_of_scope (32)
Cost saved~$8.40 (no main LLM)
→ See rejection patterns
📐 Output contract
Declare the shape the downstream code expects. Validated at write / review / publish / runtime — prevents TypeError: x.toLowerCase is not a function in <eval> nodes downstream.
Expected shape
Primitive
string — plain text
string · non-empty · ≤280 chars
boolean · "yes"/"no"/true/false
enum · one of [...]
Structured
JSON · promo.schema.json
JSON · custom schema…
JSON · array<string>
+ Define new contract…
Schema preview
{
  "headline": "string · ≤80",
  "bullets": "array<string> · ≥3",
  "cta":     "string · nullable"
}
Consumed by · 3 downstream nodes
compose_email · expects headline.toLowerCase()
compose_landing_html · iterates bullets[]
moderate_image · null-checks cta
On violation at runtime
Warn only (log)
Coerce (best-effort: object → JSON.stringify)
Block + retry once + fallback model
Block + halt run (MULTI_LAYER default)
Last 7d violations3
Recovered by retry2
Halted runs1 · #4712 (MCT key error)
→ See violation in Inspector
🚦
Quality Gate — like unit tests for prompts. A library of "must-pass" cases (tone, format, length, JSON schema, LLM-judged rubrics). Any prompt change that breaks them is blocked from shipping.
why? ▾
📋 What it does
1.Each prompt has an Eval Suite — input + expected output + criteria.
2.Criteria types: CONTAINS, REGEX, JSON_SCHEMA, MAX_LENGTH, LLM_JUDGE (with calibration).
3.Auto-runs on every prompt edit; blocks publish if regressions exceed threshold.
4.Auto-captures bad prod cases as new tests (closes the feedback loop).
🎯 Result
A safety net of regression tests that grows over time. Pass-rate sparkline shows quality maturity. Blocked publish dialog explains exactly which cases failed and why.
🧩 Process gaps closed
·No more "we tested it on one example, looked fine".
·No more silent regressions on edge cases (tone, schema, length).
·Knowledge of "what good output looks like" no longer trapped in one engineer's head.
·Eval discipline becomes part of the publish flow, not optional.
⚠ Risks mitigated
·Customer-facing regression slipping through before launch.
·"Who is liable?" — bypass requires a written reason in audit log.
·Brand violations from informal AI tone (LLM-judge on rubric).
·Schema breakage downstream (JSON_SCHEMA criterion catches it pre-prod).

🚦 Quality Gate Braintrust

5 eval suites · 92% avg pass rate · 1 suite blocking publish (tone_classifier · 87% ▼)

Eval suites · 5
promo_copy_quality
47 cases · 92% pass
json_extractor_check
120 · 100%
tone_classifier ⚠
30 · 87% ▼
summarizer_length
18 · 100%
compose_email_safety
22 · 95%
Cases · drag to reorder priority
⋮⋮acme-q2-launchP1
⋮⋮b2b-fintech-toneP1
⋮⋮short-tweet-280P2
⋮⋮json-strict-schemaP1
⋮⋮brand-voice-warmP2
⋮⋮multilingual-esP3
⋮⋮edge-empty-briefP3
⋮⋮compliance-strictP1
⋮⋮cta-includeP2
⋮⋮persona-cfoP2
+ 37 more
Critical Cases · 47
CRITICAL P1 must-pass12
REGRESSION from prod incidents15
EDGE stress / unusual8
SAMPLE happy path12
Dataset
Versioned · Apr 19 by Jess · 47 cases
Total cases
47
Pass rate
92%
Avg cost / case
$0.018
Suite cost / run
$0.85
Last run
2h ago
Pass rate history
30 days ago today
Case: b2b-fintech-tone
Inputs
{{brief}}
{{brand}}
{{tone}}
Reference output (optional)
Leave empty if relying on criteria only.
Criteria · 5
TypeConfigWeightLast result
OUTPUT_CONTRACTJSON · promo.schema.json · runtime: block+retry5✗ 2/47 type-mismatch
CONTAINS"Acme"1✓ pass
MAX_LENGTH280 chars1✓ pass
JSON_SCHEMApromo.schema.json2✓ pass
LLM_JUDGE"tone matches professional+warm rubric"3✗ 3/10
OUTPUT_CONTRACT is the strongest gate: validates type/shape of every output, including headline.toLowerCase()-style downstream consumption. Failures here block publish regardless of pass-rate (would have caught MCT key-error TypeError before reaching prod).
📊 Coverage report
Are we testing the right inputs, not just any inputs?6 dimensions · AI suggests cases for top gaps
73%
overall coverage · 3 gaps
Input length
92%
Customer segment
88%
Edge conditions
65%
Language
32%
Schema variants
28%
FB themes
12%
Coverage policy: warn at <70% · block publish at <50% (suite settings).
Top 3 gaps · AI-suggested cases
⚠ 0 multilingual cases
12 prod runs in ES this week · 0 in suite
⚠ "tone too informal" untested
Customer Feedback theme has 8 occurrences · 0 cases
↗ View feedback & capture
⚠ schema variant `cta=null`
Seen in 7% of prod outputs · no eval case
Last suite run
6
improved
3
regressed
38
unchanged
Cost delta+12%
Latency delta−8%
vs baselinerev 141 (main)
⚠ Above 5% regression threshold → Publish gate would block this change.
LLM-judge calibration
tone judge (Haiku) r = 0.84
Calibrated on 50 human-labeled cases. Threshold: r ≥ 0.7
Auto-capture from prod

Convert any production run into an eval case from Inspector or directly from Customer Feedback (👎 with comment → REGRESSION case).

3 cases captured this week from production runs.
Cassette & Replay

Record real production LLM calls into immutable cassettes; replay against any prompt branch / model swap to compare outputs without spending $$ on real API calls.

Active recording
acme-q2-prod847 calls · 2.1 MB
bravo-spring-baseline312 calls · 0.9 MB
tone-classifier-incident47 rejected · 0.2 MB
Replay scenarios
Cheap-model-swap — replay acme-q2-prod on Haiku · cost & quality delta
Prompt-fix-validation — replay tone-classifier-incident on rev 143 · check 47 fail-cases pass
Provider-switch — replay all on OpenAI · acceptance/cost diff
Cassettes are immutable, hashed, retention 90d (Pro) / 1y (Enterprise). Use in CI for prompt-change validation without real API spend.
Suite verdict
⚠ Above regression threshold — publish blocked
Pass rate 92% (43/47) on this draft vs 96% on main.
3 hard regressions: tone judge, schema, length.
If you publish anyway, an audit log entry with your reason will be created and visible to OWNER.
↗ Fix in Studio
📚
Golden Cases — upload brand guides, tone-of-voice docs, SKU catalogs once. The AI looks up the relevant snippets per request instead of you copy-pasting them into 40 prompts. Less tokens, lower cost, single source of truth, instant updates everywhere.
why? ▾
📋 What it does
1.Drop in PDFs / docs / Markdown — gets chunked & indexed automatically.
2.The kb_lookup node retrieves top-K relevant chunks per query.
3.Test queries in the playground to see what the AI will actually see.
4.Audit which chunk fed which run — full traceability.
🎯 Result
Prompts that don't repeat the same 3,200 tokens of brand context every call. Update the source PDF once → 40 campaigns pick it up automatically. Composed prompt is reproducible.
🧩 Process gaps closed
·No more "I edited the brand guide but old prompts still have the old version inline".
·No more giant prompts approaching context limits.
·No more inconsistency between campaigns about what the brand voice is.
·Brand & legal teams own the source of truth — content team just uses it.
⚠ Risks mitigated
·Off-brand AI output because someone forgot to update one prompt.
·Bill shock from giant repeated prompts (KB cuts 47%+ tokens).
·Out-of-date catalog data leaking into customer-facing copy.
·Hallucinations — AI answers from your facts, not its imagination.

📚 Golden Cases

3 collections · 15.7K chunks indexed · 2,140 retrievals this week · $0.42 embed cost MTD

Golden Cases collections · 3
📚
Brand Guides
12 docs · 1.2K chunks
📚
Product SKUs
4 docs · 8.4K chunks
📚
Past Campaigns
87 docs · 6.1K chunks
Documents in Brand Guides
📄Brand_book_v3.pdf
4.2 MB · 287 chunks · indexed 2h ago
📄Tone_of_voice.md
23 GC · 18 chunks · indexed 2h ago
📄Visual_guidelines.pdf
12 MB · 156 chunks · indexed 1d ago
📄Q1_recap.docx
⏳ indexing… 1.1 MB · 47%
📄Old_archive.pdf
⚠ outdated · 31d ago
+ 7 more
Drop PDF / MD / TXT to upload
or browse · Supported: PDF, DOCX, MD, TXT, CSV (max 50 MB)
Documents
12
Chunks
1.2K
Used by
8
CC-items
Retrievals 7d
2,140
Embed cost MTD
$0.42
Test query · debug retrieval without running a graph Ragie Pin good queries as eval cases for the RAG layer
Retrieved chunks · 3 Ragie in 1.2s · embedding cost $0.0001
#1 · 0.91 Brand_book_v3.pdf · p.14 · "Tone of voice"
"When addressing B2B audiences in regulated industries, lead with precision. Avoid colloquialisms, but stay warm. Use industry vocabulary sparingly — assume the reader is busy and skeptical."
#2 · 0.83 Tone_of_voice.md · "Audience guidelines"
"Fintech requires precision over flair. Avoid metaphors that could confuse compliance audiences. Emphasize measurable benefits, audit-friendliness, regulatory alignment."
#3 · 0.74 Brand_book_v3.pdf · p.22
"For senior decision-makers, lead with outcomes (ROI, risk reduction) before features. One-page summaries trump deep dives."
KB settings · Brand Guides Ragie
Embedding model Ragie
Changing this triggers full re-index (~$0.024).
Chunking Ragie
Retrieval Ragie
How chunks land in production prompts
Composed prompt for generate_promo_text
You are a copywriter for Acme Bank.
Tone: professional, warm.

Use insights from the Golden Cases library:
[KB chunks · 3 retrieved · 1,247 tokens]
  ─ "When addressing B2B audiences in regulated industries…" (Brand_book_v3 p.14, 0.91)
  ─ "Fintech requires precision over flair…" (Tone_of_voice.md, 0.83)
  ─ "For senior decision-makers, lead with outcomes…" (Brand_book_v3 p.22, 0.74)

Brief: Q2 launch for Acme Bank's new B2B fintech product…
Without KB: the copywriter prompt would have to inline these guidelines (~3,200 tokens of static brand-context per call). With KB: only relevant chunks (~1,247 tokens), and updates to the source PDF propagate automatically.
🧬 Vector store · Ragie
847
documents · across 12 partitions
Last sync2m ago
Indexing modehi_res
Tenant isolationpartition = companyId ✓
Dedup strategyexternal_id = fileUrl
Failed uploads (7d)3 · all retried
Powered by Ragie — handles chunking, embedding, multi-format extraction (PDF/DOCX/images with layout-awareness). We own the schema, governance, and grounding signals on top.
→ See chunks consumed by which prompts (Impact Graph)
Used by · 8 CC-items
kb_lookup_tone
Acme Q2 · last 14m
kb_lookup_brand
Bravo Spring · 1h
kb_lookup_voice
Charlie Promo · 3h
kb_lookup_compliance
Echo Black Friday · 5h
+ 4 more
kb_lookup node config (preview)
Golden Cases
Query template
Top K
Score ≥
Max tokens
Filters
brand = {{client_brand}}
lang = en
Output format
🧪
Hypothesis — turn business hunches ("a more conversational tone will work better for B2B") into proper experiments. Chat the AI to design it, run on real traffic, get a yes/no verdict with statistical evidence — instead of arguing in Slack.
why? ▾
📋 What it does
1.Frame a hypothesis: statement, expected outcome, risk, population, intervention, metric.
2.AI chat designs the experiment — type, split, sample size, guardrails, cost estimate.
3.Pre-registered analysis locks the decision rule before the run.
4.Live evidence with primary + guardrail metrics, sample outputs, AI summary.
5.Adopt → auto-merges; Reject → archived with notes; both feed institutional memory.
🎯 Result
An auditable yes/no verdict backed by statistical evidence and qualitative samples. Adopted hypotheses become production changes; rejected ones become organizational learning ("we already tested this, here's why it failed").
🧩 Process gaps closed
·No more "let's just try it on prod" experiments without controls.
·No more p-hacking after the fact (analysis is pre-registered).
·No more re-running ideas that already failed (related/conflicting card).
·Business decisions backed by data, not by who argues loudest.
⚠ Risks mitigated
·Adopting a "feels good" change that actually hurts conversion.
·Killing a viable change because someone misread early noisy data.
·Customer-facing impact during experiment (guardrails auto-pause).
·Forgetting why a decision was made 6 months later.

🧪 Hypothesis

2 hypotheses running · 2 designed · 1 drafting · 8 adopted / 23 total (35%)

All
Drafting
Designed
Running
Verdict ready
Adopted
Rejected
Common hypotheses
Tone change → acceptance ratio
Cheaper model → quality drop
Prompt cache → latency
KB enabled → consistency
Add CTA → conversion
Browse marketplace…
Pipeline · 8
▶ RUNNING · 2
Switch promo to conversational tone
A/B 50/50 · 3 days remaining
acceptancecost
Haiku on extract_intent
Shadow · 18h elapsed
costquality
📋 DESIGNED · 2
Add CTA in compose_email
awaiting OWNER approval
Brand GC → Bravo Spring
awaiting OWNER approval
✏️ DRAFTING · 1
Reduce summarizer length
Jess · started 12m ago
✓ ADOPTED · 2
Multi-shot json_extractor
+18% schema pass · adopted Apr 12
Brand context → KB
−47% input tokens · adopted Apr 5
✗ REJECTED · 1
Drop tone_classifier
−9pp acceptance · rejected Apr 14
Hypothesis quality
Adopted / total8 / 23
Avg evidence cost$24
Avg time-to-verdict3.2d
✨ Quick actions
Day 2/5 evidence trending toward ADOPT · CI [+2.1, +10.5]pp · all 3 guardrails healthy. Statistical power already 62%. You can stop early with current evidence (saves $4 + 3 days), or extend the full 5-day window for robustness.
📋 Adopt & open Change Plan
Hypothesis canvas
▶ RUNNING · day 2/5
📝 Statement
📈 Expected outcome
⚠ Risk if wrong
🎯 Population (who)
Acme Q2 Bravo Spring +1
Filter: industry=fintech
🔧 Intervention (what)
Branch promo_copywriter@conversational
vs main rev 142 (control)
📏 Metric (how measured)
Primary: acceptance_ratio
Guard: refusal rate, brand_safety, cost/run
🤖 AI chat · experiment design draft updated 2m ago
Recommended design
Type
A/B test
vs shadow / replay
Split
50 / 50
randomised by run-id
Sample size
~480 runs
to detect 5pp at p=0.05
Duration
5 days
at current 100 runs/day
Why these metrics
  • acceptance_ratio — best aligned with your hypothesis
  • refusal rate — guardrail (rises if tone too informal for compliance)
  • brand_safety LLM-judge — guardrail (catches off-brand outputs)
  • cost/run — guardrail (informal tone often → shorter outputs → minor cost change)
Estimated cost & constraints
  • Run cost: ~$8.60 (480 runs × $0.018)
  • LLM-judge eval: ~$2.40 (Haiku on 480 outputs)
  • Total: ~$11 for verdict
  • Compliance OK — same model, prompt change only
⚠ Chat warnings · 2
  • Population is not balanced: Acme Q2 has 3× more traffic than Bravo Spring → consider stratified split or per-campaign analysis.
  • Last similar hypothesis ("Drop tone_classifier") was rejected −9pp 5 days ago. Reviewer should check it doesn't conflict.
Pre-registered analysis (locked at start)
primary_winner: acceptance_ratio[treatment] > acceptance_ratio[control] + 0.02
                 AND p_value < 0.05
guardrails:
  - refusal_rate[treatment] - refusal_rate[control] < 0.02
  - brand_safety_judge[treatment] >= 7.0
  - cost_per_run[treatment] < cost_per_run[control] * 1.20
adoption_decision: ALL_PASS → suggest adopt; ANY_FAIL → reject and document why
Evidence · day 2/5
192 runs collected· $4.30 spent· on track
Primary metric · acceptance ratio
95% 85% 75% treatment 92.4% control 86.1%
Δ +6.3pp · p = 0.018 ✓
Guard · refusal rate
3% 0% limit 2pp
Δ +1.2pp · within ±2pp budget
Guard · brand safety judge
10 5 floor 7.0
Δ −0.3 · 8.4 / 10 — well above floor
Sample size
192 / 480
40% of target
Statistical power
62%
target 80% by day 5
Confidence interval
[+2.1, +10.5]pp
95% CI on Δ
Sample outputs · 6 random pairs (qualitative evidence)
▼ Control · "Acme Bank Q2 launch"
"Q2 brings precision to your treasury. Real-time visibility for CFO teams in regulated industries."
acceptance: ✓ · brand_safety: 8.7
▲ Treatment · "Acme Bank Q2 launch"
"Treasury teams, this one's for you. Real-time cash visibility, built for compliance — without the spreadsheets."
acceptance: ✓ · brand_safety: 8.4
▼ Control · "Bravo Spring fintech"
"Built for the precision your B2B audience demands. SOC2-ready, audit-friendly."
acceptance: ✗ (cancelled) · brand_safety: 9.1
▲ Treatment · "Bravo Spring fintech"
"You know the audit drill. We made it 10× faster — without skipping a control."
acceptance: ✓ · brand_safety: 8.2
Interim verdict · day 2/5
📈 Trending toward "adopt"
Primary
+6.3pp ✓
target +5pp
Guards
3 / 3 ✓
all within budget
p-value
0.018
target < 0.05
AI summary: Treatment outperforms control on the primary metric with statistical significance already on day 2. All guardrails are healthy. If trends hold, recommend adopting after the full 5-day window for robustness, or stop early if business urgency dictates.
🤖 Ask the data
Chat conversation
SF
you · 2h ago
I want to test if a more conversational tone works better for B2B fintech.
🤖
chat · 2h ago
Got it. To test this rigorously, I suggest A/B with 50/50 split on 3 fintech campaigns. Want me to draft the experiment?
SF
you · 2h ago
Yes, design it. What metrics do you suggest?
🤖
chat · 2h ago
Primary: acceptance_ratio (best aligned to your business goal). Guards: refusal rate, brand safety, cost. Sample size ~480 runs over 5 days for 80% power. Cost ≈ $11. Want to run it?
SF
you · 2h ago
Run it.
🤖
chat · 1m ago
Day 2: trending toward adopt. Both campaigns Acme and Bravo positive. Charlie has too small a sample yet. Updated the verdict card.
Related & conflicting
⚠ Drop tone_classifier
Rejected −9pp · Apr 14 · related quality drop
✓ Multi-shot json_extractor
Adopted +18% schema pass · Apr 12 · same prompt family
Knowledge artifacts
📓 Hypothesis log entry · auto-saved every change
📊 192 evidence runs · in Inspector
📥 Auto-captured 4 cases · in Eval Suite
🔗 Permalink · share with OWNER
🤔
Allocation Adviser — before automating a step, decide who should actually do it: a human, AI, deterministic code, or a hybrid. The Adviser scores 7 dimensions of the task and recommends the right setup with the rationale.
why? ▾
📋 What it does
1.Frame the task in 1-2 sentences (or load from a CC-item / Hypothesis canvas)
2.Score 7 dimensions: volume, determinism, reversibility, tacit knowledge, time-sensitivity, eval availability, liability
3.AI suggests verdict: HUMAN_ONLY / HUMAN_LED_AI_ASSIST / AI_LED_HUMAN_REVIEW / AI_AUTONOMOUS / DETERMINISTIC_NO_AI
4.Estimates monthly cost & throughput for each option
5.Saves decision as artefact (auditable + revisitable)
🎯 Result
A documented "build vs let-AI-do-it vs deterministic regex" decision before you sink a quarter into the wrong setup. Forces the team to think about volume + risk + reversibility instead of defaulting to "AI everything".
🧩 Process gaps closed
·"Default-to-AI" anti-pattern (using LLM for tasks regex would solve)
·"Default-to-human" anti-pattern (manual work where AI scales)
·Cross-team alignment on scope of human review
·Compliance: documented rationale per task
⚠ Risks mitigated
·Bill shock from LLM doing classifier work that Haiku does for $0.001 — or that a regex does for free
·Brand damage from AI-autonomous output that should have been human-reviewed
·Missed automation opportunity (high-volume manual work)
·Wrong human-in-loop level: too much (slow) or too little (risky)

🤔 Allocation Adviser

Decide who should do this task: human, AI, deterministic code, or hybrid · 23 decisions logged this quarter

Pre-filled task scenarios
Generate promo copy for B2B
Classify support tickets by intent
Approve a brand-safety image
Extract structured data from PDF
Compose customer-facing email
Choose model for a node
Recent decisions · 8
Generate promo copy
verdict: AI_LED_HUMAN_REVIEW · Apr 19 · Jess
live
Classify chat intent
verdict: AI_AUTONOMOUS · Apr 18 · Neil
live · Haiku
Extract JSON from brief PDF
verdict: DETERMINISTIC_NO_AI · Apr 17 · Sergii
JSONPath
Approve DALL·E images
verdict: HUMAN_ONLY · Apr 15 · Rebecca
brand-safety
Translate emails to ES
verdict: AI_AUTONOMOUS · Apr 14 · Jess
live · Sonnet
Schedule send timing
verdict: DETERMINISTIC · Apr 12 · Neil
cron rule
Tag campaigns by client
verdict: HUMAN_LED_AI_ASSIST · Apr 8 · Sergii
Decision distribution
HUMAN_ONLY 3
HUMAN_LED_AI_ASSIST 5
AI_LED_HUMAN_REVIEW 8
AI_AUTONOMOUS 4
DETERMINISTIC_NO_AI 3
Task descriptionframed by Jess · Apr 19, 14:08
Optional context: load from Studio CC-item · Hypothesis canvas
7-dimension scoring
low (1) — high (5) · drag sliders or click cells
Dimension12345What this means
Volume ~300 outputs/quarter — too much for hand-craft, too little for hardcore opt
Determinism Creative copy — many valid answers, no exact match
Reversibility Public landing page — bad output is recoverable but visible
Tacit knowledge Brand voice in KB; some judgement needed for B2B audience
Time sensitivity Daily turnaround acceptable — not real-time
Eval availability Eval suite exists (47 cases · 92% pass) — can validate AI output
Liability Brand-only liability — no regulatory or contractual claims
Composite score: 22/35 (63%)
Eval coverage: ✓ Yes (47 cases)
Volume threshold: ✓ ≥100/qtr (AI-economical)
Side-by-side comparison · per 100 outputsnumbers from your workspace history
OptionTime/outputCost/outputQuality (LLM-judge)ThroughputVerdict fit
HUMAN_ONLY (Jess hand-crafts)~25 min$8.509.4 / 103 / dayover-built
HUMAN_LED_AI_ASSIST (chat)~12 min$4.309.2 / 105 / dayoption
AI_LED_HUMAN_REVIEW (recommended)~3 min review$0.428.9 / 1030+ / day✓ best fit
AI_AUTONOMOUS (no review)0$0.428.4 / 10 ⚠unlimitedbrand risk
DETERMINISTIC_NO_AIn/a$0n/an/ainfeasible (creative)
Comparison includes Usage Hub historical data + Eval Suite quality + Inspector latency averages.
AI verdict
📌 AI_LED_HUMAN_REVIEW · "AI drafts, human approves"
Reasoning: Volume is AI-economical (300/qtr), eval suite already exists (92% pass), and quality drop from human (9.4) to AI-led (8.9) is within tolerance. Reviewer required because (a) public landing page, (b) tacit brand knowledge, (c) brand-only liability prefers a 3-min human approval over fully autonomous output. Saves ~22 min per output × 100 = 36 hours/qtr vs HUMAN_ONLY at −$808 cost. Saves ~$0 vs AI_AUTONOMOUS but reduces brand risk by 5×.
Apply to
promo_copywriter (Studio)
Risk tier
HUMAN_REVIEW (auto-set)
Resilience
Default + fallback OK
🤖 Ask the Adviser
Decision heuristics
RULE If volume < 50/quarter → bias HUMAN_LED
RULE If determinism = 5 AND task is structured → DETERMINISTIC_NO_AI
RULE If liability ≥ 4 → never AI_AUTONOMOUS
RULE If no eval suite AND verdict ≠ HUMAN → require eval before AI
CHECK If tacit knowledge ≥ 4 → ensure KB has the context
Linked artefacts
📝 Saved as decision td-2026-04-19-promo
🔗 Will appear in Weekly Review Card 7 (decisions logged)
🔗 Audit trail in Risk Classification + Reviews
📝
Prompt Review — like GitHub pull requests but for prompts. A teammate proposes a change → you see the diff + the eval suite result + comments, then approve or request changes. Production stays clean, decisions are auditable.
why? ▾
📋 What it does
1.Inbox of pending branches awaiting review.
2.Inline diff (red/green) of the prompt change.
3.Embedded eval result (pass rate, regressions, cost/latency delta).
4.Comments thread; Approve & merge with merge-preview ("what will happen").
🎯 Result
A merged branch with: auto eval re-run on prod, auto-rollback armed if quality drops >15% in 2h, full audit (who approved, why, with which eval result).
🧩 Process gaps closed
·No more direct edits to production by anyone with access.
·No more "Slack approval" — review evidence is in one place.
·Onboarding new prompt engineers — they can't break prod.
·Compliance can prove every prod change had review + eval.
⚠ Risks mitigated
·Unilateral changes by one person breaking 4 campaigns.
·Conflicting edits — branch model isolates work.
·Compliance failure ("show me who approved this") — answered in audit log.

📝 Prompt Review

2 pending review · 5 merged this week · avg wait 4h · 100% eval-passed

Pending · 2
tone_classifier @ stricter-rubric
sergii · 2h ago
Eval 100% Cost −2% Latency 0%
promo_copywriter @ experiment-tone-v2
jess · 38m ago
Eval 92% 3 regressions Cost +12%
Recently merged · 5
summarizer @ shorter
1d ago
json_extractor @ retry-fix
3d ago
brief_extractor @ ml
5d ago
Reviewer reliability · 30d
jess
0.9214 PRs
sergii
0.789 PRs · 1 regress
neil
0.547 PRs · 3 regress
Reliability = (no-regression merges) × (diligence score). Below 0.65 → workspace policy requires a 2nd reviewer.

tone_classifier @ stricter-rubric

in review

"Tightened the tone classification thresholds to reduce false positives. New rubric explicitly penalizes informal language for B2B contexts."

Author sergii@workspace · Opened 2h ago · Branch stricter-rubric
🔒 seal valid · approved by jess 2h ago 📋 CP-101 (approved) 🕸 17 downstream · this PR is in-scope for the approved plan
✨ Quick actions
All Quality Gate checks pass (30/30 cases · cost −2% · latency 0%) and 2 reviewers requested changes already addressed. Safe to approve & merge. Auto-rollback will arm for 2h to catch any production regression.
Prompt diff · main vs stricter-rubric
3Classify the tone of the following text:
4{{text}}
5
−6Choose: professional / casual / urgent / friendly
+6Choose: professional / casual / urgent / friendly
+7Strict rubric: text MUST contain at least 3 indicators
+8of the chosen tone (e.g. for "professional": industry
+9vocabulary, complete sentences, no contractions).
+10If unsure, return "uncertain" — do NOT guess.
11
12Output JSON: {"tone": "...", "confidence": 0.0-1.0}
Eval suite results · tone_classifier (30 cases)
8
improved
0
regressed
22
unchanged
All 30 cases pass. False positives reduced from 6 to 1 on the regression set. Model output now uses "uncertain" in 4 ambiguous cases vs forcing a guess.
Comments · 2
M
jess@workspace · 1h ago
Great change. Could you add 5 more cases for the "uncertain" path before we ship?
I
sergii@workspace · 30m ago
Done — added 5 cases to the suite. Pass rate stays at 100%.
🤖 AI Second Opinion · adversarial reviewruns on every PR · prevents reviewer from rubber-stamping a passing eval
Risk score
0.61
high · 3 concerns (1 contract)
Eval coverage
73%
missing edge cases
Behavioral diff
12%
on 100 cassette samples
⚠ Concern 1 · "uncertain" path lacks production samples
The new "uncertain" output triggers in 4 eval cases but has 0 production occurrences in the last 30d cassette. Real-world distribution may differ — consider a shadow run before merge.
→ Run cassette replay to estimate frequency
🚫 Concern 2 · output contract not exercised on branch
Contract declares headline: string · ≤80 but eval cases all stay under 60 chars. Branch wording "return options" increases array-shape risk → would trigger TypeError: x.toLowerCase in compose_email. 3 of 100 cassette samples returned an array on this branch (0 on main).
→ Add boundary-shape cases (auto-suggested)
⚠ Concern 3 · brand-voice cases not exercised
Eval suite covers tone classification but has only 2 cases tagged brand_voice. Acme Q2 brand-voice incidents (last week) suggest this dimension needs ≥8 cases to be representative.
→ Open Coverage report
✓ Validated · no regression on the 47 rejected runs from yesterday
All 47 cassette samples from the tone_classifier incident now pass with this branch. This is meaningful evidence the change addresses the root cause.
Before you approve
📋 Required diligenceREVIEW tier · 3/4 boxes
Diligence score: 0.78
workspace median 0.71 · top 0.92
→ Policy in Risk Classification
Merge preview · on Approve
tone_classifier updated to rev 167 (was 166)
Affects: Acme Q2, Bravo Spring, Charlie Holiday
Eval re-run on prod · 30/30 pass · auto-rollback armed (−15% / 2h)
Branch stricter-rubric closed · sergii notified
Audit: "Merged with 2 comments, eval 100%, expected −5 false positives"
💳
AI Usage Hub — see exactly where every dollar of your LLM bill goes (down to one node). Set budgets and hard stops. AI Advisor finds concrete savings ("switch this to Haiku → save $78/mo"). Bring Your Own Key for enterprise.
why? ▾
📋 What it does
1.Real-time spend per workspace / team / cost-center / campaign / node / model.
2.Budgets & alerts — soft warning at 50%/80%, hard stop at 100%.
3.Usage Advisor: AI suggests model swaps, prompt caching, deterministic post-processing.
4.Allocation — chargeback to clients via tags, BYOK to put your own key in path.
🎯 Result
A predictable AI bill with no surprises, plus a concrete savings projection ($164/mo → $1,968/year in this workspace) you can act on incrementally.
🧩 Process gaps closed
·No more "the AI bill exploded, why?" — root cause is one click away.
·Finance can chargeback to clients accurately (showback PDF).
·Optimization stops being a project and becomes a habit (advisor every day).
·Procurement can ask "BYOK?" and get yes.
⚠ Risks mitigated
·Bill shock — runaway costs from a misconfigured loop or batch job.
·Margin erosion — paying for premium models on tasks Haiku handles fine.
·Lost client trust on chargeback — numbers are wrong / opaque.
·Compliance / data-residency — BYOK keeps data in client's contract.

💳 AI Usage Hub Grafana Cloud Metrics

$1,247 of $3,000 spent (42%) · forecast EoM $2,840 · advisor found $164/mo savings

Apr 2026 (current)
Mar 2026
Feb 2026
Q1 2026
YTD
Custom range…
CSV (line items)
CSV (aggregated)
Showback PDF (per cost center)
QuickBooks export
Workspace budget
April 2026
$1,247 / $3,000
42% used · 11 days left
⚠ Forecast: $2,840 (95% of cap)
Per-team budgets
Marketing Acme$524 / $1,500
Marketing Bravo$398 / $1,000
Internal experiments$325 / $500
Cost centers (chargeback)
client:Acme$314
client:Bravo$201
client:Charlie$142
internal$98
✨ Quick actions
Forecast $2,840 = 95% of $3K cap (will hit by Apr 28). Usage Advisor found $164/mo savings across 3 changes — applying just the top one (Haiku swap on extract_intent) brings forecast safely under cap. Eval-gated, reversible.
Spent MTD
$1,247
▲ 18% vs Mar
Forecast EoM
$2,840
95% of cap
Avg cost / run
$0.66
▼ 8%
Avg cost / call
$0.018
stable
Daily spend · last 30 days
Usage Daily budget ($100) Anomaly
$160 $120 $80 $40 Mar 21 Apr 5 Today
👥 Team budgets & cost centers
Team / cost-centerOwnerBudgetSpent MTDUsedForecastHard stopNotify
Marketing Acme client:Acme Rebecca $1,500 $524
35%
$1,420 ✓ at 100% Slack #acme
Marketing Bravo client:Bravo Rebecca $1,000 $398
40%
$952 ✓ at 100% Email
Internal experiments internal Neil $500 $325
65%
$580 ✓ at 100% Slack #ai-internal
Content team Charlie client:Charlie Jess $300 $142
47%
$305 ⚠ at 95% (rare) Email
Echo Black Friday client:Echo Sergii $200 $198
99% ⚠
over cap ❌ stops at 100% PD + Slack
Workspace total: $1,587 spent of $3,500 in team budgets · $123 not allocated to any team
How chargeback works: tags on runs (client:, project:, team:) automatically attribute LLM cost to the right cost center. Team owner is notified at 50/80/100% of budget. Hard-stop at 100% means new LLM calls are blocked workspace-wide for that tag (in-flight runs complete).
💡 Usage Advisor — 3 findings · est. savings $164/mo
Switch extract_intent from Opus to Haiku → save ~$78/mo
It's a binary classification — Haiku scored 96% on the eval suite vs Opus's 98%. Quality delta within tolerance.
Confidence: High · Eval coverage: 30 cases · Affects: 3 campaigns
Enable Anthropic prompt caching on compose_landing_html → save ~$54/mo
87% of input tokens are static system prompt (1,820 tokens). Cache hit rate would be ~85%.
Confidence: Very High · Risk: None (caching is invisible to output)
Replace LLM-postprocessing in summarize_brief with regex → save ~$32/mo
The post-processing step extracts a JSON field — JSONPath would be deterministic and free.
Confidence: Very High · Bonus: latency drops from 1.4s to <10ms
Anomalies (24h) Grafana Cloud
Spike on Apr 14: $158
3.4× rolling avg ($46). Drove by Echo Black Friday batch run (5,000 elements).
Pre-run approvals (today)

Runs above $50 require owner approval before start.

Pending1
Approved today3
BYOK (Bring Your Own Key)
Use your provider contract

When enabled, your Anthropic / OpenAI keys will be used. Your invoice from the provider; we are no longer in the billing path.

Cost alerts
50% of monthly capnotify
80% of monthly capnotify + alert
100% of monthly caphard stop
Daily anomaly > 2×notify
🎯
Risk Classification — every CC-item has a risk tier. The platform applies different guardrails depending on the tier (review required, multi-validation, human-in-the-loop). One unified policy, no ad-hoc rules.
why? ▾
📋 What it does
1.4 risk tiers: TRUST / REVIEW / CO-CREATION / MULTI-LAYER
2.Each CC-item is classified — defaults from the type
3.Higher-tier nodes get stricter Quality Gate, mandatory human approval, multi-judge eval, redaction enforcement
4.Audit log of misclassification attempts (e.g. demote DALL·E to TRUST blocked)
🎯 Result
Brand-risky outputs (DALL·E images, public emails) cannot ship without human approval regardless of who edits the workflow. Compliance has a single answer to "what controls are in place per output type?".
🧩 Process gaps closed
·"oops we made the AI auto-publish to Twitter" can't happen
·Compliance audit: matrix of risk tier × controls
·Reduces accidental over-trust of LLM in critical paths
⚠ Risks mitigated
·Reputational: off-brand image goes public without review
·Legal: customer-facing email contains hallucinated facts
·Regulatory: compliance can't prove human-in-the-loop

🎯 Risk Classification

71 CC-items classified · 4 MULTI_LAYER · 18 HUMAN_REVIEW · 7 CO_CREATION · 42 TRUST

All
TRUST
REVIEW
CO-CREATION
MULTI-LAYER
Risk tiers · 4
TRUST_BY_DEFAULT
Reversible · low impact · auto-publish OK
42 CC-items
HUMAN_REVIEW
Reviewer must approve before publish
18 CC-items
CO_CREATION
Human-in-the-loop iteration · chat-style
7 CC-items
MULTI_LAYER
Eval + judge + human approval + audit
4 CC-items
Misclassification attempts
2026-04-18 — Jess tried to demote compose_email to TRUST · blocked
2026-04-12 — Neil queued DALL·E for auto-publish · blocked

TRUST_BY_DEFAULT

42 CC-items · default for most types

Outputs are reversible, low-impact, audience small or internal. Publishing is automatic. Failures don't damage brand.

Default for
prompt, json_prompt, summarizer, classifier
Required gates
Eval Suite (≥1 case)
SLA on incidents
Best-effort
Controls matrix · who applies what
ControlTRUSTREVIEWCO-CREATIONMULTI-LAYER
Eval Suite required≥1 case≥10 cases≥10 + LLM-judge≥30 + 2 judges
Pre-publish gateauto-passblock on regressionblock + human OKblock + 2 humans
Reviewer required1 reviewer1 reviewer + author2 reviewers (incl. compliance)
Hallucination detectoroffwarn-onlywarn-onlyblock on suspect
PII redactionpassivepassiveactiveactive + audit log
Resilience policyanyanyrequire exact model OKRequire Exact Model
Fallback to other modelallowedallowed (logged)warn userblocked
Audit retention30 days90 days1 year7 years
Customer report inclusionaggregatedper-run availableper-run + transcriptper-run + judge rationale
Output contract (type/shape)recommendedrequired · warn-onlyrequired · block + retryrequired · block + halt + audit
Artifact immutability after approvalnone · agent free-editseal-on-merge · agent edit allowed but flagged in auditseal-on-merge · agent edit auto-breaks seal → re-review requiredseal-on-merge · any edit (human or agent) requires Change Plan + 2nd reviewer
Intent allowlist (Guardrails)offallowlist + warnallowlist + refuseallowlist + refuse + audit
Context grounding floorwarn if <0.6require ≥0.7 or refresh KBblock if <0.8
Vector-store availability (Ragie)graceful · skip retrieval, log warninggraceful · alert · run completes without RAG contextblock if RAG unavailable · queue + retryblock + halt + page on-call · no degraded retrieval
Eval coverage floor≥60%≥75%≥90% across all dimensions
Reviewer diligence friction2/4 mandatory checks3/4 mandatory + AI second-opinion4/4 + AI second-opinion + adversarial cassette
Reviewer reliability floor≥0.65 or 2nd reviewer auto-added≥0.80 or compliance reviewer added
Anti-rubber-stamp + intent + grounding controls (added 2026-04) — see Prompt Review for the per-PR enforcement, and Quality Insights for the Grounding monitor.
CC-items in this tier · 42🤔 Allocation Adviser
CC-itemTypeCampaignAuto-classified byLast edit
extract_intentjson_promptAcme Q2type-default2h ago
summarize_briefpromptAcme Q2 · Bravotype-default1d ago
tone_classifierjson_promptshared (3)type-default14:10 today (Jess)
fetch_briefhttpAcme Q2type-default3d ago
post_processfunctionAcme Q2type-default5d ago
+ 37 more
High-risk items (MULTI_LAYER)
📧 compose_landing_html (Acme)
Customer-facing landing page · public URL · brand-risk
2 reviewers 7yr audit Block hallucination
🖼 generate_dalle_hero
Auto-generated brand image · staging-bucket only · moderation required
2 reviewers PII redaction Brand-safety judge
📨 send_email_blast
Sends to 50K+ recipients · irreversible · CAN-SPAM
🌐 publish_to_microsite
Public S3 deploy · no rollback after CDN
Compliance posture
SOC 2 — separation of duties✓ enforced
GDPR Art.22 — human in loop✓ enforced
ISO/IEC 42001 — AI controls⚠ partial
EU AI Act — transparency✓ enforced
🕸
Impact Graph — answers "if I change this prompt / contract / KB chunk, what breaks?" in one click. The artifact graph (prompts ↔ DAG nodes ↔ eval suites ↔ contracts ↔ campaigns ↔ feedback) made first-class and queryable.
why? ▾
📋 What it does
1.Builds a graph of every artifact: prompts, DAG nodes, eval cases, output contracts, KB chunks, campaigns, feedback themes.
2.Shows blast radius for any change — direct + transitive consumers.
3.Pre-computes verdict per consumer: green / will re-eval / will need re-review / will break.
4.Acts as router — every other tool can ask "give me everything downstream of X".
🎯 Result
Before any prompt edit, you see exactly which 4 campaigns / 18 eval cases / 2 contracts will be touched. No more "ship and pray". Especially valuable for MCT — one upstream prompt fans out to 12 client campaigns.
🧩 Process gaps closed
·"Invisible dependencies" — transitive impact you wouldn't see otherwise.
·MCT replication risk — see all client tenants affected by one shared prompt.
·Contract change without warning to downstream <eval> consumers.
·Stale KB chunk that 7 prompts silently rely on.
⚠ Risks mitigated
·Cross-campaign breakage from a "small" prompt tweak.
·Forgotten contract consumer crashes after schema change.
·One MCT edit silently degrades all replicated client microsites.

🕸 Impact Graph

Selected: tone_classifier · 17 downstream artifacts · 3 will break · 4 campaigns · 8 eval cases · 2 contracts · 1 MCT

downstream (consumers)
upstream (sources)
both
1 hop (direct)
2 hops
3 hops
unlimited
📋 Open as Change Plan →
Pick an artifact
Prompts · 14
tone_classifier
17 downstream
promo_copywriter
11 downstream
extract_intent
6 downstream
compose_email
3 downstream
compose_landing_html
2 downstream
Output contracts · 9
promo.schema.json
8 consumers · 2 nodes
tone_result.contract
5 consumers
Golden Cases · KB Ragie · 6
brand_guides (Acme) · 142 chunks
stale 6w · 14 consumers via 3 retrieval nodes
compliance_rules_eu · 89 chunks
stale 8w · 9 consumers via 2 retrieval nodes
Retrieval ccItems · 17 Ragie
retrieve_brand_voice
topK=8 · feeds 4 prompts
retrieve_compliance
topK=12 · feeds 2 prompts
MCT (shared) · 4
mct_b2b_finance
12 client tenants
🚫 Pending edit on tone_classifier would break 3 downstream consumers
Jess opened branch experiment-tone-v2. Impact analysis shows: 1 contract violation (tone_result.contract shape would change), 1 MCT (mct_b2b_finance replicates to 12 client tenants), 1 stale eval suite (tone_judge_v2 uses brand_guides rev 7, source updated 6w ago). Open as Change Plan to capture mitigation steps before merge.
📋 Open as Change Plan 📝 Open PR with impact attached
Blast radius · 3 hops · 17 artifactsedges show data flow direction (left → right)
SELECTED DIRECT (1 hop) TRANSITIVE (2 hops) SURFACE (3 hops) 📝 tone_classifier prompt · rev 142 → 143 (draft) 📐 tone_result.contract 🚫 will break · shape change 🚦 tone_judge eval suite ⚠ will re-run · 30 cases 📦 mct_b2b_finance 🚫 replicates to 12 tenants 🔀 promo_dag · node 4 ⚠ requires re-review 🔀 chat_assistant_dag ✓ unaffected 📋 brand_voice eval (4 cases) ⚠ coverage drops to 65% ⚙ <eval> compose_email 🚫 .toLowerCase risk 📊 Quality Gate suite ✓ auto-runs 🏢 12 client tenants 🚫 next MCT sync breaks 📝 Reviews queue ⚠ 2 PRs need re-review 📋 Coverage report ⚠ brand_voice 50% → 30% 📚 brand_guides (Acme) ⚠ stale 6w · likely re-sync 📢 Acme Q2 🚫 publish would block 📢 Bravo Spring 🚫 publish would block 📢 Charlie Holiday ⚠ degraded acceptance 📢 Golf B2B ⚠ acceptance drift 💬 Customer Feedback ✓ pattern detector armed
Will break (3) · solid red edge
Re-eval / re-review (8) · solid amber
Unaffected (4) · dashed green
17 nodes · 18 edges · max depth 3
🔎 Killer query · ask the graph anythingnatural-language → graph query
Other useful queries: "show me all prompts that depend on stale KB chunks" · "which contracts have no eval coverage on edge cases" · "every artifact a single reviewer can break alone" · "MCTs whose downstream tenants have a pending PR"
Verdict per consumer
🚫 3 will break
  • tone_result.contract — output shape change
  • mct_b2b_finance — 12 client tenants downstream
  • compose_email.toLowerCase() on changed shape
⚠ 8 will need re-review / re-run
  • • 2 eval suites · auto re-run on merge
  • • 2 PRs in queue · new diligence required
  • • 4 campaigns · expect acceptance drift 5-15pp
✓ 4 unaffected
  • chat_assistant_dag
  • • Quality Gate suite · normal trigger
  • • Customer Feedback — pattern detector ready
  • • Weekly Review · cards refresh on Sun
Affected reviewers · 4
jess (tone_classifier owner)notif: ✉
neil (mct_b2b_finance)notif: ✉ + slack
sergii (compose_email)notif: ✉
rebecca (Acme account)notif: ✉ (FYI)
Graph health (workspace)
Total artifacts847
Edges2,141
Orphans (no consumer)12
Stale-KB consumers38
MCTs (replicated)4 · 47 tenants
RAG retrieval nodes17 · feed 23 prompts
Last graph rebuild14:48 (2m ago)
Edge types
CONSUMES1,124
GATES312
EXERCISES287
RETRIEVES_FROM186 · Ragie
REPLICATES_TO142 · MCT
FED_BY90
📋
Change Plan — review the intent to change before any artifact is generated. Author proposes a structured plan (which prompts, contracts, eval cases will move), human signs off the plan, then Studio opens for actual edits. Prevents "ship and discover".
why? ▾
📋 What it does
1.Author drafts a structured Change Plan — scope, motivation, expected impact.
2.Auto-fills affected artifacts from Impact Graph (pre-computed blast radius).
3.Reviewer approves / requests edits to the plan — not the code.
4.Approved plan unlocks Studio + Workflow Editor edits; only the listed artifacts can be touched.
🎯 Result
Architecture decisions are reviewed before tactical implementation. Reviewer-burnout from "approving big diffs after the fact" goes down. Approved plan becomes part of the audit trail (compliance-friendly).
🧩 Process gaps closed
·"Why are we changing this?" — answered before the diff exists.
·Scope creep mid-PR ("while I was in there, I also fixed…").
·Architectural debate happens at 4-line plan stage, not at 200-line PR review.
⚠ Risks mitigated
·Major refactor sneaking in under cover of a "small fix" PR.
·Two authors silently editing the same artifact.
·Reviewer rubber-stamping a 500-line diff because they trust the author.

📋 Change Plan

3 active · 1 awaiting approval · 12 merged this week · avg plan→merge cycle 2.4d

Active plans · 3
CP-104 · tone tightening
jess · awaiting approval · 38m ago
3 will break 12 MCT tenants
CP-103 · Haiku swap
neil · approved · in Studio
approved cost −$78/mo
CP-102 · grounding contract
sergii · in PR review
in implementation
Recently merged · 5
CP-101 · intent allowlist
2d ago
CP-098 · brand_guides re-sync
4d ago
CP-095 · output contract on promo
6d ago

CP-104 · tone tightening

awaiting approval

Author jess@workspace · Created 38m ago · Reviewer neil (auto-routed: tone_classifier owner)

📐 Plan structureedit fields directly · saved as draft
Motivation (why now)
Scope · what will change
Expected impact (auto-filled from Impact Graph)
🚫 3 will break
tone_result.contract · mct_b2b_finance · compose_email
⚠ 8 re-review
2 eval suites · 2 PRs · 4 campaigns
✓ 4 unaffected
chat_dag · QG suite · feedback · ritual
→ View live in Impact Graph
Mitigation steps before merge
  • Coordinate with neil (mct_b2b_finance owner) — schedule MCT replication for off-peak window
  • Add Output Contract migration note: tone_result.contract v2 (backwards-compatible coercion for 14 days)
  • Re-fresh brand_guides KB chunk before publish (currently stale 6w)
  • Notify Rebecca (Acme account) of expected acceptance drift 5-15pp on Day 1
Success criteria (measurable)
  • • Acceptance ratio recovers to ≥85% within 2h post-merge
  • • Zero new OUTPUT_CONTRACT_VIOLATION events in 24h
  • • Acme CFO confirms tone fix on next-day campaign
Rollback condition
🤖 AI plan critiquelooks for missing mitigations, hidden scope, alternatives
⚠ Missing in scope
Plan touches tone_result.contract shape, but doesn't list the contract migration deliverable explicitly. Recommend adding "+ Bump contract version to v2" to Scope.
💡 Alternative considered
Could add a downstream coercion node instead of breaking the contract. Pros: zero MCT impact. Cons: technical debt; hides the fix from compose_email author. Recommend: stick with current approach but document the trade-off here.
📚 Past similar plan
CP-076 (3 months ago, same author) tightened summarizer rubric — similar shape, took 4 days from plan→merge with 1 contract bump. This plan is on a similar trajectory.
Review checklist
📋 Required for approvaltier: HUMAN_REVIEW
After approval
jess gets Studio edit access on listed artifacts only
Any out-of-scope edit triggers a "Plan amendment required" prompt
Draft PR is auto-created with the plan attached as PR description
4 affected reviewers get FYI notification ("CP-104 plan approved, PR coming")
Approval signature stored in audit log (compliance-grade)
Plan history
14:38Plan created (jess)
14:42Impact auto-attached
14:51AI critique posted (3 items)
15:02Mitigation owners contacted
15:14Reviewer assigned (neil)
Awaiting approval
📅
Weekly AI Quality Review — every Monday Neil + Jess + Rebecca walk through the same 7-card agenda: alerts, drift, regressions, hypothesis status, advisor wins, GC updates, action items. AI prepares the deck.
why? ▾
📋 What it does
1.Auto-prepares Monday 10am deck from the week\'s data
2.7-card agenda: alerts, drift, regressions, hypothesis, savings, KB, action items
3.Each card has owner + decision pre-suggested by AI
4.Action items become tickets with owners + deadlines
🎯 Result
30-min weekly meeting that keeps quality on a steady-state curve · institutional memory of "what we tried, what worked, what we decided".
🧩 Process gaps closed
·No more "we should review quality... at some point"
·Single source of cross-functional alignment
·Action items don\'t fall through the cracks
⚠ Risks mitigated
·Quality drift goes unnoticed week-over-week
·Knowledge silos (PROMPT-only, OPS-only, biz-only)
·Customer reports compiled in panic instead of routinely

📅 Weekly AI Quality Review Grafana Cloud Reporting

This Monday Apr 19, 10:00 UTC · 7-card agenda · 4 attendees · 4 action items pending

Apr 14-19 (this week)
Apr 7-12
Mar 31-Apr 5
Past 12 weeks…
This week\'s agenda · 7 cards
1 · Alerts & incidents2
2 · Quality drift3
3 · Regressions caught5
4 · Hypothesis status2 running
5 · Usage Advisor wins$78/mo
6 · GC updates1
7 · Action items4
Previous reviews
Apr 7-12
3 actions · all closed
Mar 31-Apr 5
5 actions · 4 closed
Mar 24-29
2 actions · all closed
Attendees
NNeil (OWNER)
JJess (PROMPT)
RRebecca (Account)
SSergii (PROMPT)
⚠ Card 1 · Alerts & incidents (2)
Apr 19, 14:48 — Acceptance ↓ 17pp on tone_classifier
Cause: Jess direct-edit on main · no eval ran. Resolved 14:55 by Neil rollback. 47 rejected runs. Acme account notified by Rebecca.
resolved <30 min5 cases captured to suite
Decision needed: should direct edits on main be disabled for shared prompts?
Apr 19, 12:42 — Anthropic 5xx burst (36 min)
Auto-fallback to OpenAI · 47 fallbacks · +$8.40 cost · 0 user-facing failures.
no human action needed
📊 Card 2 · Quality drift signals (3)
PromptDirectionΔOwnerAction
tone_classifier↓ acceptance−17pp (resolved)Jess5 eval cases added
json_extractor↑ refusal rate+5× over weekSergiiinvestigate provider policy update
compose_email↑ length+18% avg tokensJesstighten output spec
✓ Card 3 · Regressions caught (5 last week)
5/5 blocked at Quality Gate before reaching production
By prompt: tone_classifier (2), promo_copywriter (2), summarizer (1)
By author: Jess (3), Sergii (1), Rebecca (1)
Estimated incidents avoided: ~3-4 (extrapolated from severity)
🧪 Card 4 · Hypothesis status (2 running)
"Conversational tone for B2B fintech" (Rebecca · day 2/5)
Trending ADOPT · CI [+2.1, +10.5]pp · Acme +9.1pp · Bravo +4.8pp · Charlie −1.2 (small n)
Decision suggested: extend Charlie monitoring 3 days, adopt for fintech subset
"Haiku on extract_intent" (Neil · shadow A/B 18h)
96% pass · within ±2pp · save $78/mo at full rollout
Decision suggested: ADOPT after 24h window completes
💰 Card 5 · Usage Advisor wins · $78/mo realized · $86/mo queued
extract_intent → Haiku
$78/mo · in shadow A/B
compose_landing_html cache
$54/mo · queued for next week
summarize_brief regex
$32/mo · queued
📚 Card 6 · GC updates (1)
Tone_of_voice_v4.md uploaded by Brand team Apr 19, 13:55
18 chunks · old v3 archived · 4 KB-using campaigns automatically picked up · acceptance stable, no semantic break.
📋 Card 7 · Action items (4) — assigned during today\'s meeting
ActionOwnerDueLinked to
Disable direct edits on shared prompts (require branch)NeilApr 24Card 1
Investigate json_extractor refusal spikeSergiiApr 22Card 2
Decide on Hypothesis #4 adoption for CharlieRebeccaApr 22Card 4
Schedule compose_landing_html cache rolloutNeilApr 26Card 5
Rolling 12-week trend
Acceptance trend86% → 87.3%
Avg incidents/week1.8
Avg regressions caught4.2
Avg savings/week$42
Action close rate92%
Ritual health
Last skippedMar 17 (1 mo ago)
Avg duration28 min
Action items / meeting4.1
Ritual is healthy — no missed weeks in 4 weeks, all actions closed within deadline.
Auto-prepare deck

Every Monday 09:30 UTC the system auto-builds this deck from:

  • Quality Insights · alerts & drift
  • Reviews · merged + bypassed
  • Hypothesis · running & adopted
  • Usage Advisor · realized + queued
  • KB · updates & impact
  • Tickets & action items closed since last week

🔀 Business flows

13 end-to-end flows · sequence diagrams of who triggers what, who responds, where decisions happen

All flows
Investigation only
Authoring only
Automated
Customer-facing
📚 What this page is
Each flow is a UML-style sequence diagram: actors are shown as vertical lanes (personas + features + external systems), time flows top to bottom, and arrows show messages between them. The icon and color of the arrow tells you the type of interaction.
user action AI/system action alert/event success/result decision/warning self-call (feature acts on itself)
🔄 How flows influence each other
Flow 1 → Flow 2: incident response auto-captures cases into Eval Suite → next prompt change in Flow 2 catches the same regression at publish-time.
Flow 2 → Flow 7: every approved review in Flow 2 becomes a line item in the customer quality PDF (Flow 7).
Flow 3 → Flow 5: cost-bounded retries from Reliability (Flow 5) feed Usage Hub forecast in Flow 3.
Flow 4 → Flow 2: adopted hypotheses become canonical prompt branches; subsequent edits in Studio (Flow 2) use them as baseline.
Flow 5 → Flow 1: if fallback degrades quality (rare), Quality Insights catches the dip and triggers Flow 1 — incident response.
Flow 6 → Flow 2: GC updates can shift retrieval semantics; Flow 2 picks up the change because Eval Gate catches breakage automatically.
Flow 7 ← all: customer report rolls up evidence from Flows 1-6.
Flow 8 → Flow 1: Customer Feedback (Flow 8) detects pattern before quality alert fires (Flow 1) — closes the loop end-to-end. Inspector run back-link makes the diagnosis instant.
Flow 8 → Flow 4: high-volume feedback themes seed Hypothesis ideas ("informal tone consistently rejected — should we test conservative variant?").
Flow 9 → Flow 1+2+3: a properly-scaffolded campaign (Workflow Editor + risk + budget + eval) reduces incident probability by ~10× vs. ad-hoc setup. Allocation Adviser verdict prevents both default-to-AI and default-to-human anti-patterns.
Flow 10 ← Flow 4: Hypothesis Lab is the substrate for Model Migration — same A/B mechanics, just treatment = different model. Cassette & Replay makes it cost-free.
Flow 10 → Flow 3: successful migration → Usage Advisor records the savings; future deprecations follow same playbook.
Flow 11 → Flow 8: Intent Classifier rejections aggregate as themes in Customer Feedback inbox · 47 "legal_advice" rejects in a week → product seeds a new product surface (Hypothesis Lab in Flow 4).
Flow 11 → Flow 3: rejected-before-LLM requests show up in Usage Hub as cost-prevented signal; the cheap-classifier-as-gate pattern becomes a workspace default.
Flow 12 → Flow 1: AI second-opinion catches what eval missed → an incident that would have triggered Flow 1 next day never happens. Diligence-score policy makes this systematic, not heroic.
Flow 12 ← Coverage Report: coverage gaps surfaced during PR review feed straight back into the eval suite — quality grows over time without anyone scheduling a "quality sprint".
Flow 13 → Flow 2 + 12: Plan-first changes shift the review center of gravity from PR-time (Flow 2/12) to plan-time. Flows 2 and 12 still happen, but the architectural debate is already settled — they become tactical safety nets, not the only line of defense.
Flow 13 prevents Flow 1: a 12-tenant MCT incident that would have triggered Flow 1 (incident response) the next day never happens, because Impact Graph surfaced the blast radius before any code was written.
Flow 4 → Flow 13: adopted hypotheses now have a "📋 Adopt & open Change Plan" CTA — the experiment-design discipline of Flow 4 carries straight through to the planned-change discipline of Flow 13.
Flow 13 ← all artifact graphs: every feature that produces edges in the Customer Context Graph (Studio, Workflow Editor, Quality Gate, Output Contract, KB, MCT, Reviews) feeds Impact Graph and therefore Flow 13.

🗺 Feature map

19 features grouped by lifecycle stage · for the demo conversation with the customer

📚 What this page is
Every feature: why it exists, the business problem it solves, the concrete result, key arguments for customer conversations, and how features compose. Grouped by where they sit in the workflow lifecycle: Hub · Build · Lab · Operate · Settings.
🧭 Hub
🏗 Build · creation tools — author, test, review & ship the AI stack
🔀
BUILD new
Workflow Editor
Visual DAG of CC-items. The actual no-code pipeline builder customers use.
Problem: graph errors at runtime; non-engineers can't safely connect prompts to KBs to guards.
Result: shippable campaign DAG · type-checked I/O · cost & latency known before first run · graph-level linter blocks dangerous topologies.
Args: 14+ CC-types · per-node deep-link to Studio/Eval/Risk · linter (cycles, MULTI_LAYER without human_approval) · graph-level versioning · cost preview per run.
✏️
BUILD
Prompt Studio PromptLayer
IDE for prompts: branches, A/B playground, inline linter, cost preview, Memory tab, Risk tab.
Problem: "edit prod and see what happens"; non-engineers can't author safely.
Result: tested-before-shipping prompt with cost & latency known · reviewer signs off · 1-click rollback.
Args: 5 model cost previews · drag variable chips · linter · diff vs main · publish triggers Quality Gate · Memory tab · Allocation Adviser inline.
📐 Output Contract (new): declare expected shape (string / JSON-schema / array) on Risk tab · enforced write-time (linter) → review-time (AI second-opinion) → publish-time (Quality Gate) → runtime (Inspector OUTPUT_CONTRACT_OK/VIOLATION events). Catches TypeError: x.toLowerCase in <eval> nodes before MCT-key bugs reach prod.
📝
BUILD
Prompt Review
GitHub-style PRs for prompt changes: diff + eval result + comments + merge preview.
Problem: Slack approvals · evidence scattered · no compliance trail.
Result: single approval queue with diff + eval + cost/latency delta · auto-rollback armed · audit-ready.
Args: red/green inline diff · embedded eval inline · merge preview lists affected campaigns · pulls from Studio, blocks via Quality Gate.
🤖 AI Second Opinion + Diligence (new): adversarial pre-merge review surfaces concerns the eval missed · mandatory 4-box checklist with auto-fill from behavior signals · per-reviewer reliability score · friction policy by risk tier (anti-rubber-stamp).
📚
BUILD
Golden Cases
Upload brand guides once. RAG retrieval into kb_lookup nodes. Single source of truth for AI context.
Problem: brand guides copy-pasted into every prompt → bill shock, inconsistency.
Result: update one PDF → 40 campaigns pick it up · ~47% input-token reduction.
Args: drag-and-drop upload · 3 chunking strategies Ragie · reranker + hybrid search Ragie · test-query playground · audit retrieval · directly drives Usage Advisor savings.
🎯
BUILD
Risk Classification
Per-CC-item risk tier. Platform applies different controls per tier.
Problem: high-risk outputs (DALL·E to public S3) get same thin pipeline as internal classifier.
Result: brand-risky outputs cannot ship without human approval · controls matrix per tier · compliance answer in 1 minute.
Args: 4 fixed tiers · controls matrix 9×4 · misclassification audit · compliance posture (SOC2/GDPR/ISO 42001/EU AI Act) · enforced by Quality Gate, Reviews, Reliability.
🛡 Intent Validation (new): per-CC-item allowlist of intents · classifier refuses out-of-scope queries before they hit the expensive LLM · refusal templates · weekly stats on rejected requests + cost saved.
🔒 Artifact immutability (new): approved prompts/contracts/eval-cases are sealed (hash + reviewer + timestamp) · any agent or human edit auto-breaks the seal and triggers re-review at tier-appropriate strength.
🚦
BUILD
Quality Gate Braintrust
Test suites for prompts. Block regressions; auto-capture bad prod cases. Includes Cassette & Replay for cost-free validation.
Problem: "tested it on one example, looked fine" → silent regression.
Result: growing safety net · pre-publish dialog explains failures · bypass requires written reason.
Args: 6 criterion types incl. LLM_JUDGE with calibration · Cassette & Replay (record real prod, replay against branch/model) · auto-capture from Inspector + Customer Feedback · Critical/Regression/Edge/Sample tags.
📊 Eval Coverage report (new): 6-dimension coverage matrix (length, segment, edge, language, schema, FB themes) · AI suggests cases for top gaps · "70% pass on 4 cases" anti-pattern detector.
📋
BUILD new
Change Plan
Review the intent to change before any artifact is generated. Author drafts a plan, reviewer signs off the plan, then Studio unlocks for actual edits.
Problem: reviewers approve 200-line diffs because they trust the author · scope creep mid-PR · architectural debate happens too late.
Result: motivation/scope/expected-impact/mitigation discussed at 4-line plan stage · approved plan unlocks Studio (only listed artifacts editable) · audit-grade signature.
Args: structured plan (motivation, scope, impact, mitigation, success criteria, rollback) · auto-fills affected artifacts from Impact Graph · AI plan critique · approval signature stored in audit log · scope guard prevents out-of-list edits.
🧪 Lab · answers under uncertainty — design experiments & decisions
📡 Operate · daily ops & observability — get information, take action
🔍
OPERATE · FOUNDATION
Inspector Langfuse
Permanent record of every run with the actual prompt & response. The backbone every other feature builds on.
Problem: "we can't reproduce" — engineering bottleneck for every quality complaint.
Result: any complaint resolved <5 min · final artifacts · audit trail for SOC2/GDPR · production cases pinnable to eval suite in 1 click.
Args: per-LLM-call payload · resolved variables · retry/fallback markers · permalinks · 3-tier retention · PII redaction · feeds Quality Insights, Quality Gate, Usage, Hypothesis evidence, Customer Feedback back-refs.
🕸
OPERATE new
Impact Graph
Live graph of every artifact (prompts ↔ DAG ↔ eval ↔ contract ↔ KB ↔ campaigns ↔ feedback). Killer query: "if I change X, what breaks?"
Problem: "small" prompt edit silently breaks 12 MCT-replicated tenants · contracts have invisible downstream consumers · stale KB chunks support 7 prompts no-one tracks.
Result: blast-radius preview before any merge · per-consumer verdict (will break / re-review / unaffected) · auto-notification of affected reviewers · feeds Change Plan + Reviews.
Args: 847 nodes · 2,141 edges · 3-hop depth default · natural-language queries · Cypher export · graph rebuilt every 2-5 min · MCT replication-aware (one upstream prompt → N tenants visible).
📈
OPERATE
Quality Insights
Tracks acceptance, drift, alerts. Hallucination detector + ChatOverride accept-rate dashboard.
Problem: nobody knows when AI quality is degrading until customers complain weeks later.
Result: alerts in minutes · recovery plan with cases auto-captured.
Args: primary + guardrail metrics · auto-correlation with prompt edits · Hallucination Detector · ChatOverride dashboard · feeds Hypothesis baselines.
🧭 Context Grounding Monitor (new): 6-signal score (KB freshness, retrieval, drift, memory loss, budget, var resolution) · stale-KB offender list · 4.2× correlation with hallucinations.
💬
OPERATE new
Customer Feedback
Explicit 👍/👎 + comments from end-customers. Triage queue, NPS/CSAT, pattern detection.
Problem: only implicit signal — actual recipient voice never reaches the iteration backlog.
Result: direct line from end-customer to eval suite · pattern detection ("3rd 'tone too informal' on rev 142") · NPS + CSAT.
Args: embed widget (1-line snippet) · sentiment + theme tagging · per-campaign breakdown · Slack notifications on negatives · auto-link to Inspector run that generated the output.
💳
OPERATE · FINOPS
Usage
Visibility · Control · Optimization · Allocation. Per-team budgets, hard stops, AI Usage Advisor, BYOK.
Problem: AI bill explodes; nobody knows why. Finance can't chargeback to clients.
Result: predictable bill · Advisor recommendations · per-team accountability · client showback PDF.
Args: per-team budgets with hard-stop · forecast vs cap · Usage Advisor (3 actionable findings) · prompt cache · BYOK · cost-center attribution.
📅
OPERATE
Weekly Review
Monday 30-min ritual: 7-card auto-prepared deck (alerts, drift, regressions, hypothesis, savings, KB, action items).
Problem: "we should review quality… at some point". Action items fall through cracks.
Result: 30-min weekly meeting keeps quality on steady-state · institutional memory.
Args: auto-prepare Sunday 22:00 · 7 fixed cards · action items with owners + deadlines · 12-week rolling trends · push to ticketing.
⚙ Settings · workspace configuration
🔗 How features compose — concrete dependency map
Inspector is the foundation. Every feature consumes its persisted run + LLM-call data: Quality Insights plots, Hypothesis evidence, Usage per-call cost, Quality Gate auto-capture, Customer Feedback back-refs, Cassette recording.
Workflow Editor binds prompts together. Per-node deep-link to Studio. Graph publish triggers all attached eval suites. Studio "Used by" reverse-shows which graphs use a prompt.
Golden Cases feed Studio. kb_lookup variables become draggable chips. Reduces token cost — reflected in Usage Advisor.
Studio feeds Reviews + Quality Gate. Branch → Request Review → Reviews queue. Click Publish → Quality Gate eval suite (incl. Cassette replays) → blocks if regressions.
Quality Insights feeds Hypothesis & Quality Gate. Drift identifies underperforming prompts → Hypothesis candidates. Captured cases → eval tests.
Customer Feedback closes the loop end-to-end. 👎 + comment → pattern detection → "Add to Eval" or "Open Hypothesis". Inspector run back-link shows which output triggered.
Cassette & Replay enables cheap iteration. Real prod calls recorded → replayed against branches/models without API spend. Hypothesis verdicts snapshot cassette for future replay.
Reliability events surface in Inspector + Usage. Each fallback / retry logged with marker; cost overhead tagged.
Risk Classification gates everything. Quality Gate, Reviews, Reliability, Hallucination Detector all consume the resolved tier.
Allocation Adviser sets the tier. Before authoring: decide AI/human/deterministic. Verdict auto-sets risk tier + creates linked artifacts.
Hypothesis is the decision layer. Reads from Inspector, uses Studio (treatment branch), Cassette (snapshots), Quality Gate (auto-capture), Usage (cost). Adopted hypothesis → Reviews merge.
Model Catalog is the substrate. Defines what models exist; Studio model-picker, Reliability fallback chain, Usage spend-by-model all read from catalog.
Team Budgets enforce per-team accountability. Tag-based attribution from Inspector calls; hard-stop gates LLM calls when team hits 100%.
Weekly Review aggregates everything. Sunday 22:00 auto-prep pulls from Quality Insights, Reviews, Hypothesis, Usage Advisor, Customer Feedback, KB updates.
Campaign Explorer aggregates ALL. Recommended-action banner = correlation engine over alerts + edits + spend + feedback + reviews + hypothesis status.
▶ Demo flow (12 min with the client) — full end-to-end
  1. Campaign Explorer → "Acme Q2" pinned card → recommended action banner.
  2. Quality Insights → 17pp drop · likely cause (prompt edit) · Hallucination Detector card.
  3. Customer Feedback → Acme CFO complaint matches the pattern — "3rd this week".
  4. Inspector → run #4828 with the actual bad output.
  5. Workflow Editor → see the prompt's position in the DAG · linter · cost preview.
  6. Prompt Studio → fix → Publish → Quality Gate blocks regression with Cassette replay validation.
  7. Prompt Review → teammate change with merge preview.
  8. Always-On AI → policy editor + simulator.
  9. Model Catalog → BYOK + Opus sunset planning.
  10. Golden Cases → Test query for retrieval.
  11. Hypothesis → "tone → conversational" interim verdict.
  12. Usage → Team budgets + Advisor $164/mo · "Apply with eval check".

🏷 Vendor landscape — what we'd build vs. what we'd buy

9 categories of AI infrastructure · 22 vendors surveyed · 5 marked as drop-in for our prototype today

📚 What this page is for
A walk through the AI-infrastructure market with one question for each component: "is this still our differentiation, or has it become a commodity service we should buy instead of build?" Drop-in vendors are flagged with a ↪ vendor pill across the prototype — this page is the master list with the rationale behind each call.
Sources: vendor websites, Crunchbase / PitchBook funding data (early 2026), public customer logos. Tier judgement is opinion, not endorsement.
📌 Summary — what to flag in the prototype today
✓ Already integrated
  • Ragie — RAG / vector store (Chris's branch)
🟡 Flagged as drop-in candidates
  • Langfuse → Inspector
  • Braintrust → Quality Gate
  • Portkey → Always-On AI
  • PromptLayer → Prompt Studio
🟢 Build (still our differentiation)
  • • Output Contract / Intent Validation
  • • Hypothesis Lab
  • • Usage Hub (cost optimization layer)
  • • AI Second Opinion / Coverage Report
  • • Impact Graph / Change Plan / Artifact Seal
  • • Risk Classification matrix
The asymmetry: 5 commodity layers can be bought (saves engineering quarters); the rest is what makes VelocityEngine recognizable to a customer. The vendor pills in the prototype make this trade-off legible to anyone reviewing the design — including you, when scoping next sprints.
📊 Grafana Cloud · pills across the prototype
Blocks tagged with a Grafana Cloud pill are realizable on our existing Grafana Cloud subscription without building new infra. Green-coloured Grafana Cloud · live pills mark what's already shipped.
✓ Live today
  • Quality Insights — structured-logging dashboard at veprod.grafana.net
🟡 Realizable next
  • Inspector → Cloud Traces (Tempo, OTel LLM)
  • Always-On AI → Cloud SLO · OnCall · Incident
  • Usage Hub → Cloud Metrics + anomalies
  • Customer Feedback → NPS/CSAT panels
  • Activity timeline → Cloud Logs (Loki)
  • Weekly Review → Cloud Reporting (PDF export)
Why this matters: Grafana Cloud covers the operational infra (traces, metrics, logs, alerts, on-call, SLOs, incidents) — freeing us to focus engineering on the campaign-domain layer on top.
🧭 How to read the "buy vs. build" call
Maturity tier
Funded Series A+ or community-OSS-standard, multi-tenant, public API + docs, real customer logos. We do not bet on toys.
Replaces real engineering
Genuine months-of-work substitute, not a thin wrapper. Building this internally would compete with the vendor's full-time team.
Customer can tell the difference
If a customer can distinguish "this is the vendor part" from "this is your part", flag it. If your part is just glue, you are not differentiated.
1 · Retrieval & Vector Store already integrated
What it doeschunking · embedding · partition-based multi-tenant retrieval · multi-format extraction (PDF/DOCX)
Vendors at our tier
Ragiein use
Managed RAG with hi_res layout-aware extraction; partition-based isolation; reranker + hybrid search.
Pineconeprimitive · DIY chunking
Vector DB only — you own chunking and embedding pipeline. Series B+, enterprise grade.
Weaviateprimitive · OSS+cloud
OSS vector store with managed offering; same DIY trade-off as Pinecone.
What WE still own
  • Source-of-truth governance: Golden Cases collections, freshness-vs-source tracking, KB drift monitoring.
  • Grounding score: 6-signal cross-check (KB freshness, retrieval relevance, drift, memory loss, budget, var resolution).
  • Stale-KB consumer alerts: "38 prompts depend on a chunk that hasn't been re-indexed in 6 weeks".
  • Workflow integration: retrieval as first-class CC-item with type-checked I/O.
  • Audit grade: per-run RAG_QUERY events with chunk-level provenance for compliance.
Verdict: ✅ Buy. Done. Chris has the integration in branch RAG-ccItem.
2 · LLM Observability / Tracing flagged in prototype
What it doesper-call payload · resolved variables · cost · latency · retry/fallback markers · session grouping
Vendors at our tier
Langfuserecommended
OSS + cloud; ~10k GitHub stars, $4M YC W23. Samsara, Khan Academy on logos. De-facto OSS standard.
HeliconeYC W23 · proxy-based
Proxy-style observability + caching. 2k+ companies. Lower friction; higher per-call latency overhead.
Arize Phoenix / AXenterprise · $70M Series B
OTel-native; Uber, ServiceNow. Heavyweight — overkill if you only need per-call inspector.
What WE still own (Inspector)
  • Business-context tagging: campaign ID, asset type, MCT lineage. Vendor sees calls; we see which campaign for which client.
  • Domain events: SEAL_VERIFIED, OUTPUT_CONTRACT_OK, INTENT_VALIDATED, RAG_UPLOAD_FAILED, ARTIFACT_SEAL_BROKEN. These are our policy primitives, not observability.
  • "Reproduce in <5 min" UX: Inspector → Studio → Eval Gate → Customer Feedback links — that's a product, not a trace viewer.
  • Audit-grade retention: 3-tier retention with PII redaction tied to risk classification.
Verdict: 🟡 Strongest "buy" candidate. Frees engineering to focus on the campaign-context layer.
3 · Eval-as-a-Service flagged in prototype
What it doesdatasets · LLM-judge calibration · regression detection · run history · CI integration
Vendors at our tier
Braintrustrecommended
Series A $36M (a16z); Notion, Stripe, Airtable, Zapier. The vocabulary closest to our Quality Gate.
Langfuse Evalsbundled with tracing
Evals + dataset versioning + judge templates inside the same product as observability. One-vendor story is attractive.
Patronus AI$17M seed · pre-built evaluators
Generic evaluators (hallucination, PII, toxicity); good complement, not a Quality Gate replacement.
What WE still own (Quality Gate)
  • Campaign-quality rubric: brand voice, promo correctness, regulated-industry tone — vendors don't know what "good copy for Acme Bank" looks like.
  • Coverage Report: 6-dimension matrix that tells where the suite is weak, not just which cases pass.
  • Auto-capture from Inspector + Customer Feedback: the loop from real complaint → permanent regression test.
  • Publish-gate enforcement: the gate sits in the Studio publish flow with seal-and-revert mechanics. Vendor would be the eval engine, we own the policy.
Verdict: 🟡 Strong "buy" — eval orchestration is commodity now.
4 · Prompt Management flagged in prototype
What it doesversioning · branches · A/B · labels (prod/staging) · rollback · non-engineer editing
Vendors at our tier
PromptLayerrecommended
Production at ENS, Gorgias, Speak. Established 2022; non-engineer editing UI is mature.
Langfuse Promptsbundled · linked to traces
Versioned prompts integrated with their tracing. Single-vendor consolidation play.
LatitudeOSS · younger
OSS prompt-engineering platform with versioning + evals. Smaller ecosystem.
What WE still own (Studio)
  • Risk tab: Output Contract, Intent Validation, runtime policy declared per CC-item.
  • Change-Plan-gated edits: can't open Studio for production prompts without an approved plan in scope.
  • Impact Graph reverse-deps: "this edit will break 12 MCT tenants downstream" — vendor doesn't see your DAG.
  • Linter with domain rules: contract violation risk, role-tagging on user data, KB usage hygiene.
  • Multi-stakeholder review workflow: AI Second Opinion + diligence checklist + reviewer reliability score.
Verdict: 🟡 "Buy" the storage / versioning layer; keep our authoring UX wrapping it.
5 · LLM Gateway / Failover flagged in prototype
What it doesmulti-provider routing · retry · circuit breaker · cost cap · semantic caching
Vendors at our tier
Portkeyrecommended
$3M seed; production at Postman, Springworks. 200+ models; cleanest gateway UX.
OpenRouter~$100M ARR · routing-only
Massive dev adoption; pure routing without observability/policy depth.
LiteLLMOSS standard · $6M
OSS proxy + hosted; 100+ providers, virtual keys, budgets. Adobe, Lemonade as users.
What WE still own (Always-On AI)
  • Per-CC-item policy: "MULTI_LAYER prompts must use claude-opus-4.7, no fallback"; vendor handles routing, we set the policy.
  • Tier-aware enforcement: Risk Classification matrix decides what fallback is allowed for each tier.
  • Compliance guards: Require Exact Model + region restrictions tied to client contracts.
  • Reliability event correlation: fallback events surface in Inspector with cost overhead tagged.
Verdict: 🟡 "Buy" the plumbing; keep tier-aware policy as our governance layer.
6 · Guardrails / Output Validation partially mature — careful
What it doesJSON schema · PII · topic filtering · structured output · hallucination check
Vendors at our tier
Guardrails AI$7.5M seed · OSS+cloud
Hub of validators; ~4k stars. Generic validators, not domain rules.
Patronus AI$17M seed
Runtime guardrails (hallucination, retrieval relevance, custom policies); enterprise focus.
NeMo GuardrailsNVIDIA · self-host OSS
Colang DSL for dialog flow; not SaaS — self-host.
Why we did NOT flag this
  • • JSON schema validation is commodity (Pydantic / Instructor cover it natively).
  • Output Contract is more than schema — it's a 4-phase lifecycle (write / review / publish / runtime) tied to risk tier and downstream consumer analysis. Vendors don't model that.
  • Intent Validation is a cheap-classifier-as-gate pattern that lives before the LLM call; vendors validate after.
  • • Domain rules (no competitor names, no hallucinated SKUs) require workspace-specific policies vendors don't host.
Verdict: 🔴 Build. The category is real but immature; "drop-in" framing would mislead the customer.
7 · Feature Flags & Experimentation flag platform + our glue
What it doesfeature flags · A/B · stats engine · sequential testing · power calculations
Vendors at our tier
Statsig$100M Series C · OpenAI customer
Ex-Facebook experimentation team; built-in stats; AI experiment templates.
LaunchDarklypublic-tier · enterprise-priced
Mature feature flag platform with experimentation. Heavyweight for AI cadence.
Eppo$20M Series A · warehouse-native
DraftKings, Perplexity, Cameo. Stats over your data warehouse.
Why we did NOT flag this
  • • No vendor is purpose-built for "LLM Hypothesis Lab" — they're general A/B platforms.
  • PICO-style framing + AI-designed experiment + pre-registered analysis is our authoring layer; vendor would only run the math.
  • • Cassette-replay-based experiments (no real prod cost) is unique to our stack — vendors don't model it.
  • • Adopted-hypothesis → Change Plan transition is workflow logic, not statistics.
Verdict: 🟢 Build now; consider Statsig later as the stats engine if experiment volume grows.
8 · LLM FinOps / Cost Observability no clear winner yet
What it doesper-call cost · per-customer/team rollups · budgets · anomaly detection · forecast
Vendors at our tier
Heliconeprimitive · per-call
Per-call cost attribution + budgets. Useful primitive, not a finance product.
Langfuse Costsprimitive · trace-based
Cost per trace/user/session; alerts. Built for engineering, not FinOps.
Vantage / CloudZeroFinOps + nascent LLM SKU
Mature FinOps; LLM-specific support is just emerging.
Why we did NOT flag this
  • • No "Datadog for LLM cost" winner exists yet; everyone builds on top of Helicone/Langfuse primitives.
  • Usage Advisor ("switch this prompt to Haiku → save $78/mo") is a domain-aware optimizer — vendors give numbers, we give recommendations.
  • Per-team budgets + chargeback PDF + BYOK live in our workspace model; vendors are blind to it.
  • Cost-bounded retries tied to risk tier — that's gateway + policy, not a cost dashboard.
Verdict: 🟢 Build. Re-evaluate in 12-18 months.
9 · Red-teaming / Adversarial Testing security-focused, not quality
What it doesjailbreak detection · adversarial prompts · synthetic edge cases
Vendors at our tier
Promptfoo~5k stars · OSS+paid
Red-team plugins, adversarial datasets, CI integration. Shopify, Discord, Anthropic-internal usage.
Lakera$10M seed · enterprise
Runtime + offline red-teaming; "Gandalf" famous; Dropbox, Citi customers.
Patronus AI$17M seed · adversarial sets
Generic red-teaming suites; complement to in-house quality coverage.
Why we did NOT flag this
  • • Maturity is real for security red-teaming; we're focused on quality coverage.
  • AI Second Opinion is workflow-aware (cassette replay + coverage report + contract risk) — vendor red-teams in isolation.
  • Coverage Report identifies which dimension is undertested; vendors generate cases without that targeting.
  • • Promptfoo is interesting as a reference for case-generation patterns; not a drop-in.
Verdict: 🟢 Build (quality-coverage); evaluate Lakera if security red-teaming becomes a sales requirement.

📖 Documentation

Product Guide · auto-generated from the codebase by deep-code-analysis

Done