mars.deck
01 / 17
SharePoint On-Prem Security · 2026

MARS

Multi-Agent Resolution System

Automated bug reproduction, fix, and verification for SharePoint On-Prem — powered by collaborative AI tools.

Presented by Prajwal Kambale · Intern – SharePoint On‑Prem team  ·  contributions: Worker tool & orchestrator

The Problem

Every security bug follows the same six-step manual workflow — a developer is pulled in at every step (highlighted), each costing hours to days.

80+
security cases pending
80+
App-sec bugs pending
Manual bug-fixing workflow — developer in the loop at every step
AI tools surface new vulnerabilities faster than the team can fix them — the backlog only grows.

The Solution — MARS

Vision: a developer pastes a bug report, walks away, and comes back to a verified fix with a PR ready to raise.

MARS pipeline — Worker/Reviewer across every checkpoint
MARS system architecture
Agent communication diagram

Before CP1 — Port the SharePoint Online Fix (optional)

Many On-Prem bugs are already fixed in SharePoint Online. When a SPO PR exists (--spo-pr <url>), MARS studies it first and reuses that work — without skipping On-Prem reproduction.

SPO PR→ Analyze (pre-CP1)→ spo-solution.md→ CP2 candidate→ CP1–CP4 still run
Worker — one analysis turn
  • Runs before CP1, worker-only — no repro, no code yet
  • Fetches the SPO PR diff (az repos pr show / Azure DevOps REST); blocked, never faked if the PR can't be read
  • Writes spo-solution.md: fix summary, changed files, security mechanism, On-Prem pros / cons, portability mapping
  • Appends one spo_port_analysis message with an initial recommendation
How it enters CP2
  • The SPO fix becomes one labelled candidate in the plan — not a replacement for CP1/CP4
  • direct_spo_port (port as-is where On-Prem allows) or hybrid_spo_plus_mars
  • Weighed against MARS's own root-cause approaches in spo-comparison.md
  • Auto-pilot picks the best candidate; else the developer chooses — SPO is never forced

On-Prem CP1→CP4 still runs in full — SPO porting only seeds the plan with a proven fix to adapt.

CP1 — Reproduce with Proof

CP1 reproduce-with-proof flowchart
Worker · VM-A
  • Connects + smoke-tests the sandbox VM and SharePoint farm
  • If the farm isn't live, the step is blocked — never faked
  • Sends the exploit payload straight at the vulnerable endpoint
  • Records evidence.json with per-step proof + a repro tier (FULL / LIMITED)
Reviewer · VM-B
  • Reads the Worker's evidence and the original work item
  • Re-runs every step from scratch on a second, independent VM
  • Approves only if the bug genuinely reproduces and is honest
  • Sends it back with findings if it's weak or a likely hallucination

CP2 — Create a Fix Plan

CP2 fix-plan flowchart
Worker
  • Traces the root cause — why it happens, not just where
  • Drafts 2–3 fix approaches, each with exact code changes
  • Optional model-diverse ensemble proposes competing plans
  • An optional SPO port enters here as a labelled candidate (direct / hybrid)
  • Picks the winner, or hands the choice to the developer → CP3
Reviewer
  • Adversarially validates each approach against the real source
  • Checks every plan is sound, complete and actually fixes the bug
  • Approves the strongest plan to move forward
  • Rejects with findings, triggering a re-plan (≤ 2 attempts)

CP3 — Write & Validate the Fix

CP3 code flowchart
Worker
  • Implements the fix in src\, behind a debug-flag kill switch
  • Builds a debug DLL (internal build (debug DLL), retries ≤ 3)
  • Deploys to the sandbox VM and re-runs the repro
  • Iterates until the bug is dead (up to 5 attempts, developer-extendable)
Reviewer
  • Re-tests the fix on its own VM, from scratch
  • Performs an independent code review of the change
  • Approves only when the bug no longer reproduces
  • Loops back to CP3 on any failure or weak fix

CP4 — Verify, Build, Ship

CP4 verify flowchart
Worker
  • Builds the full retail patch (internal build (full patch))
  • Produces the signed .msp patch packages
  • Installs it on the sandbox VM (PSConfig / farm upgrade)
  • Re-runs the repro — it must now fail, with no regressions
Reviewer
  • Installs the same patch on its own VM
  • Independently verifies the fix and regression-safety
  • Approves → CP5 once the patch is proven
  • Sends it back to CP3 if verification fails

CP5 — PR Description (worker-only)

CP5 PR-description flowchart
Worker
  • Gathers the CP1–CP4 evidence and the change diff
  • Writes pr-description.md: title, root cause, fix, testing, evidence links
  • Appends exactly one pr_description message (append-only)
  • Produces paste-ready PR text for the developer
No Reviewer
  • CP5 is not reviewed — a single Worker turn
  • No second-VM check and no approval gate
  • The developer pastes the description into the pull request
  • Still runs even when unit tests are skipped (PR metadata is produced)

CP6 — Unit Tests

CP6 unit-tests flowchart
Worker
  • Adds/updates repo-native unit tests for this bug + the CP3 change only
  • Adds regression coverage for the vulnerable input/behavior
  • Runs the suite in the repo's native framework
  • Or submits a concrete "not-applicable" reason if no test fits
Reviewer
  • Verifies tests are correctly scoped (this bug + fix only)
  • Checks framework conventions — behavior/regression focused
  • Confirms the tests actually ran, not just compiled
  • Approves → PR ready, else loops back to fix the tests
Pipeline metrics diagram

Why You Can Trust It

MARS is built to be verifiable, not just automated. Five design choices keep it honest.

Two-VM verification

The Reviewer re-runs everything from scratch on an independent VM — nothing ships on the Worker's word alone.

Evidence gates

Every checkpoint requires machine-checkable proof (evidence.json); unproven claims are blocked, never faked.

Debug-flag kill switch

Fixes land behind a flag, so a change can be toggled off instantly without a rebuild or redeploy.

Developer in control

Interrupt, redirect, or take over at any checkpoint; escalations always route back to a human.

Crash-safe by design

State lives in one append-only JSONL log — a run can resume exactly where it stopped.

Anti-hallucination

Independent re-check + evidence + confidence scoring surfaces discrepancies before they reach a PR.

"

Q & A

MARS — Multi-Agent Resolution System