feat(athena): structured synthesis header + audit cross-check flow

Add universal YAML front-matter header to all synthesis.md outputs (council,
question, date, members, session_ids, mode, intent, responded count).

AUDIT: structured finding format with required fields (severity, confidence,
members reported, evidence, impact, fix direction). New cross-check flow asks
user to verify minority/solo findings via a second council round before
proceeding to findings processing.

Also fix prompt-file cleanup: skip cleanupPromptFile when movePromptFile
already succeeded.
This commit is contained in:
ismeth
2026-03-14 20:09:28 +01:00
committed by YeonGyu-Kim
parent 974eb8e176
commit 44b3efe7a3
3 changed files with 64 additions and 13 deletions
+43 -11
View File
@@ -3,16 +3,50 @@ import { FOLLOWUP_OR_DONE_QUESTION_BODY } from "./shared-action-paths"
export const AUDIT_GUIDANCE = `
<runtime_synthesis_rules>
Use AUDIT synthesis.
- Output numbered findings grouped by confidence: unanimous, majority, minority, solo.
- For each finding: issue, impact, evidence, fix direction.
- Output numbered findings grouped by confidence: UNANIMOUS (all or nearly all members agree), MAJORITY (more than half agree), MINORITY (2-3 members), SOLO (single member).
- Each finding MUST use this exact format:
#### #{number}: {title}
- **Severity**: {Critical|High|Medium|Low}
- **Confidence**: {Unanimous|Majority|Minority|Solo} ({N} members)
- **Members Reported**: [{member1}, {member2}, ...]
- **Issue**: {description}
- **Evidence**: {file:line references, code snippets}
- **Impact**: {what breaks or degrades}
- **Fix Direction**: {concrete remediation approach}
- After all findings, include a summary table with columns: #, Finding, Severity, Agreement, Members Reported.
- End with a Priority Recommendations section grouping findings by action urgency.
- If any findings are dismissed as false positives, list them in a Dismissed section with reasoning.
</runtime_synthesis_rules>
<runtime_action_paths>
Path type: ACTIONABLE.
1) Build actionable findings list from audit synthesis.
1) After synthesis, check for minority and solo findings (findings reported by fewer than half the council members).
2) Ask how the user wants to process findings:
2) If minority or solo findings exist, ask the user about cross-checking BEFORE processing:
Question({
questions: [{
question: "The audit found {N} findings with low agreement (minority/solo). Cross-checking sends these to ALL council members for independent verification. This costs one additional council round.",
header: "Cross-check?",
options: [
{ label: "Cross-check low-confidence findings", description: "Launch a new council round where all members evaluate the {N} minority/solo findings" },
{ label: "Skip cross-check", description: "Proceed to findings processing with current confidence levels" }
],
multiple: false
}]
})
3) If user chose cross-check:
- Build a cross-check prompt listing each minority/solo finding with its full details (ID, title, severity, evidence, fix direction).
- Use the standard council flow: prepare_council_prompt(prompt=cross_check_prompt, mode=same_mode, intent="AUDIT") -> athena_council -> background_wait -> council_finalize.
- The cross-check prompt should instruct each member: "For each finding below, independently evaluate against the codebase. For each: (1) AGREE or DISAGREE, (2) your severity rating, (3) evidence supporting your assessment. Wrap your evaluation in <COUNCIL_MEMBER_RESPONSE> tags."
- After collecting cross-check results, re-synthesize: update each cross-checked finding with the new agreement level, note changed assessments, promote/demote findings based on cross-check votes.
- Write the updated synthesis.md (overwrite the previous one in the same archive_dir).
- Include a Cross-Check Results section showing how each finding's confidence changed.
4) If no minority/solo findings exist, OR user chose to skip cross-check, proceed to processing mode:
Question({
questions: [{
question: "How would you like to process the findings?",
@@ -26,12 +60,12 @@ Question({
}]
})
3) Branch by processing mode.
5) Branch by processing mode.
Mode: One by one
- Per-finding fix options MUST be derived from the synthesis's fix directions / suggested fixes.
- Each option represents a remediation APPROACH (e.g., "Add input validation", "Refactor to use parameterized queries"), NOT an execution agent.
- NEVER use agent names (Atlas, Prometheus, Hephaestus, Sisyphus) as per-finding option labels — those belong ONLY to Step 5 (execution method).
- NEVER use agent names (Atlas, Prometheus, Hephaestus, Sisyphus) as per-finding option labels — those belong ONLY in Step 7 (execution method).
- Handle variable fix counts per finding:
- If a finding has NO fix suggestion from synthesis -> show only Skip / Defer / Stop review.
- If a finding has exactly ONE fix suggestion -> show that fix + Skip / Defer / Stop review.
@@ -124,13 +158,13 @@ Question({
})
- Resolve selected buckets into concrete finding IDs.
4) If actionable findings count is 0 OR user selected no findings, do NOT ask execution action. Ask:
6) If actionable findings count is 0 OR user selected no findings, do NOT ask execution action. Ask:
Question({
questions: [{
question: "No findings were selected for action. What should we do next?",
${FOLLOWUP_OR_DONE_QUESTION_BODY}
5) If selected findings exist, ask what action to take on the selected findings:
7) If selected findings exist, ask what action to take on the selected findings:
Question({
questions: [{
question: "How should we handle the selected findings?",
@@ -140,18 +174,16 @@ Question({
{ label: "Fix now with (Hephaestus)", description: "Hand off to Hephaestus for direct implementation" },
{ label: "Fix now with (Sisyphus)", description: "Hand off to Sisyphus for collaborative implementation" },
{ label: "Create plan (Prometheus)", description: "Hand off to Prometheus for planning and phased execution" },
{ label: "Cross-check with council", description: "Launch a new council session with this synthesis as context" },
{ label: "No action", description: "Review only - no delegation" }
],
multiple: false
}]
})
6) Execute selected action:
8) Execute selected action:
- Fix now with (Atlas) -> switch_agent(agent="atlas") with ONLY selected findings
- Fix now with (Hephaestus) -> switch_agent(agent="hephaestus") with ONLY selected findings
- Fix now with (Sisyphus) -> switch_agent(agent="sisyphus") with ONLY selected findings
- Create plan (Prometheus) -> switch_agent(agent="prometheus") with ONLY selected findings
- Cross-check with council -> launch a new council session with the current synthesis as context. Restart from Step 2 (council setup) with the synthesis included in the prompt.
- No action -> acknowledge and end
</runtime_action_paths>`
+15
View File
@@ -270,6 +270,21 @@ After Step 7, you will receive a separate runtime guidance message injected by c
Treat the injected runtime guidance as authoritative over generic defaults.
Universal requirements (all intents):
- ALWAYS start synthesis.md with this YAML front-matter header:
---
council: {archive_dir basename, e.g. council-auth-review-a1b2c3d4}
question: {original user question, verbatim}
date: {ISO 8601 date of synthesis}
members: [{member1}, {member2}, ...]
session_ids: [{bg_xxx}, {bg_xxx}, ...]
mode: {Solo|Delegation}
intent: {AUDIT|PLAN|EVALUATE|DIAGNOSE|EXPLAIN|CREATE|PERSPECTIVES|FREEFORM}
responded: {N}/{total}
---
- The header values come from your council session: question and intent from council_finalize args, members and session_ids from Step 6 launch, mode from Step 2, responded count from council_finalize result.
- After the front-matter, write the intent-specific synthesis content following the runtime guidance.
- Track agreement/disagreement across members and use agreement level as a confidence signal.
- Flag single-member points as lower confidence.
- Be concrete and evidence-based.
@@ -144,8 +144,12 @@ export function createCouncilFinalize(
})
}
let promptFileMoved = false
const relPromptFile = args.prompt_file
? await movePromptFile(args.prompt_file, base, absArchiveDir, relArchiveDir)
? await movePromptFile(args.prompt_file, base, absArchiveDir, relArchiveDir).then((result) => {
promptFileMoved = true
return result
})
: undefined
const relMetaFile = join(relArchiveDir, "meta.yaml")
@@ -164,7 +168,7 @@ export function createCouncilFinalize(
const guidance = buildAthenaRuntimeGuidance(resolvedIntent, resolvedMode)
return JSON.stringify(result, null, 2) + "\n\n" + guidance
} finally {
if (args.prompt_file) {
if (args.prompt_file && !promptFileMoved) {
await cleanupPromptFile(args.prompt_file, base)
}
}