WLCG Open Technical Forum #12 · 2026‑08‑25

AI and Facilities

Towards trustworthy adoption of agentic management and use

Blueprint diagram: a systems administrator (MCP broker + Kubernetes API access) and a physics analyst (agentic analysis pipelines) each connect via USB-C-style links through an MCP broker to a WLCG AF or site's compute, GPUs, storage, and Kubernetes control plane
Slides AI-assisted by Claude (Anthropic) · title illustration AI-generated.

The argument

Facility context and trust challenges

The model layer

Moves fast. Not ours to own.

The facility layer

Context, identity, trust, packaging, equity.

Two trust problems: AI acting on the facility, and AI acting through it.

Grounding

Assumed vocabulary

Agent

An LLM loop: decide, act, observe, repeat.

MCP server

A service exposed as tools + live context.

Skill

A reusable recipe, loaded on demand.

Runtime

The agent’s hands.

skills + MCPs + agents → a plugin; plugins → a marketplace.

01

Existence proof

In production today:
systems agents & a user gateway

Agents operating the infrastructure · physicists reaching it through an MCP broker.

Existence proof

AI at every layer of the facility

Analyst-facing

Find data, generate, submit, fit.

Operations

Reports, triage, human-approved emails.

Sysadmin-facing

Privileged copilots; trust earned, gated.

Daily HTCondor cluster report posted to Slack by the AI monitor agent

448 of 14,671 held

Drafted held-job email posted to Slack, approved by a human, then sent by the agent

draft → “approve” → sent

Same context, different blast radius.

Deep-dive · one triage run

Root causes, not symptom lists

HTCondor Queue Triage — 2026-04-01
Total held: 825 / 2395 (34%) — above threshold

 #  User      Held  Cause                    Action
 1  user-01    284  PERIODIC_HOLD >7d wall   Restructure, resubmit
 2  user-02    165  errno 13: /scratch/      Fix node perms, release
                    perms c111,c113          (node bug — blameless)
 3  user-03    155  PERIODIC_HOLD >7d wall   Restructure, resubmit
 4  user-04    121  Short queue >4h          Split <4h chunks
 5  user-05     65  errno 2: .tgz missing    Investigate, resubmit

Wall-time holds: 560/825 (68%) — three users, policy issue
Priority: fix /scratch perms c111+c113, then release user-02
(the agent does NOT do this autonomously)
  • Reads HTCondor via MCP; no ssh.
  • Root causes: a blameless node bug ≠ user error.
  • Never releases jobs autonomously.

Our expert caught 2 errors in 10 minutes: AI proposes, experts decide.

The user gateway, live

“One gateway. Your entire facility.”

mcp-portal.af.uchicago.edu · live

AI Assistant
MCP Gateway
Auth Policy Audit Broker
Rucio
HTCondor
Jupyter
Storage

Credentials stay behind the broker

Raw tokens never reach the assistant.

Add tools, not complexity

Authorize every call

Group-based access; audit logging.

How it works: part 03.

mcp-portal.af.uchicago.edu · live diagram reproduced from the portal’s landing page (af-mcp-platform).

02

AI acting on the facility

The sysadmin’s view:
earning trust for privileged actions

Monitoring · privileged operations · trust principles — earned like a new operator: incrementally, observed, revocable.

Earned autonomy

A trust ladder, with promotion gates

Observe

read-only diagnostics

Suggest

a human approves

Submit job

scoped, reversible

Restart service

restarts stuck workers

Operate production

remediates, deploys

Promotion by evaluation: fault scenarios, scored, regression-tested. Autonomy measured, not assumed.

Deep-dive

Governance lives in a git repository

Principles

Auditability, minimal privilege, human-in-the-loop.

Policies

“Read-only by default;” “no outbound web.”

Runbooks: the atomic unit of trust

PR-gated: autonomous vs human-required.

governance/
├── principles/
│   ├── auditability.md
│   └── minimal-privilege.md
├── policies/
│   ├── condor/query-rate.yaml
│   └── network/egress.yaml
└── runbooks/
    ├── htcondor/queue-health.md
    └── kubernetes/cluster-health.md

AI drafts → expert reviews → merged. The PR is the audit record.

Deep-dive · blast radius

Three governance levels

Level 1 · Assisted

Humans execute. Blast radius: a document.

Level 2 · User-scope

User credentials only. A quota.

Level 3 · Privileged

Scoped accounts. The account.

Dev sandbox
Staging cluster
Production

Nothing skips a step.

Deep-dive

What the agent knows

Normative source

What it may do.

Runtime state record

What it has learned.

DR archive

Session memory

Runbook
Agent learns
Weekly review
PR

Autonomy changes are PRs, a trust escalation log: when you stopped watching.

Deep-dive · the flip side

No context? It should stay quiet

What happened

Our agents confidently recommended Lustre tuning, for a Ceph filesystem. Plausible, fluent, wrong.

The rule: ungrounded → silent. Context is what makes an agent safe to trust.

Deep-dive · enforcement

Enforcement is dual

Policy awareness inside the agent; technical enforcement outside.

Podsk8s default
Sandboxbubblewrap, sydbox
Policynode-level
Privilegeearned

The honest gap: runtimes advise; they can’t hard-deny. Blocking needs a node policy engine or sandbox.

Meme captioned I've got to keep control

Deep-dive · a security reality

Agents don’t get the open web

  • Facilities block agent web access: prompt injection.

Stale-prone

use /path/2024-egamma-recs.config

Durable

“ask chATLAS for the current path.”

Minimal durable guidance + a trusted live source: a facility decision.

03

AI acting through the facility

The physicist’s view:
laptop → facility

The physicist hardly leaves their laptop; the facility’s job is a safe port to plug an agent into — the broker · self-service access · pipelines · identity.

The USB-C port of the facility

A credential-brokered MCP gateway

One socket. The client never holds raw credentials: the broker authenticates, mints short-lived credentials, audits every call.

Claude, Gemini or any MCP client to the mcp.af.uchicago.edu front door (oauth2-proxy, AF Keycloak OIDC), through a FastMCP aggregator and AF credential broker (FastAPI /v1) doing identity, authZ, credential brokering and audit, fanning out to rucio-mcp, ami-mcp, openmagic, panda-mcp, condor-mcp, gitlab-mcp, jupyter-control and an Nth backend with no code change

Live at mcp.af.uchicago.edu: Giordon Stark’s af-mcp-platform.

af-mcp-platform · rucio-mcp · ami-mcp · atlasopenmagic · panda-mcp · condor-mcp (Bockelman, CHEP 2026).

Self-service, today

Getting connected takes one command

Any MCP client, one endpoint:

claude mcp add --transport http atlas-af \
  https://mcp.af.uchicago.edu/mcp
  • 401 → OAuth discovery → browser login. No manual tokens.
  • No OAuth flow? Mint a token at the portal.

mcp-portal.af.uchicago.edu

Link identities, mint x509/VOMS proxies, manage tokens, see your backends.

Deep-dive · the bottleneck

Who is this client, and may it act as you?

Who is it?
Do we trust it?
How much authority?

Today: DCR. Worth watching: CIMD (identity as a URL). Atop WLCG tokens + AARC.

Deep-dive · CIMD vs DCR

Client ID = a URL the server fetches

GET /authorize?client_id=https://app.example/oauth.json&…
        │  authn server fetches the JSON, validates, applies policy
        ▼
{ "client_id":"https://app.example/oauth.json",
  "client_name":"…","redirect_uris":[…] }   ← hosted by the client
  • Wins: no /register abuse; stable identity; domain allow/deny.
  • Costs: outbound fetches (SSRF); localhost ambiguity; domain ≠ trust.
Auth today reuses grid creds (x509) or Keycloak OIDC — OAuth bridge docs.

One pipeline, four personas

Agent personalities

data_agent“get the dataset”

transform_agent“make the histograms”

batch_agent“run the cluster”

inference_agent“fit the limit”

Rucio · dCache

coffea · ServiceX

HTCondor · Dask

pyhf · RooFit

Personas speak intent, never backend names: portable across facilities.

04

Beyond one facility

The WLCG question:
coordinate what, exactly?

Options

What could be common?

Standardize the port

a conformant interface profile

Authorize the clients

client conventions on WLCG tokens

Share the knowledge

runbooks and skills, single-source

Operate together

shared eval scenarios

Which, if any, do we do together?

AARC · EOSC · ESCAPE · IRIS-HEP/OSG-LHC concepts doc.

Deep-dive · the commons

One commons, not a dozen forks

  • Who hosts, who maintains, who stops the forks?
  • ATLAS today: a single-source marketplace; reference, don’t copy.
  • An eval/lint harness keeps skills orthogonal.

Open: the cross-experiment contribution. ESCAPE is one model.

US AI proto-collaboration update: DPF 2026.

Deep-dive · ESCAPE

The agentic layer, meet ESCAPE

  • Data Lake (Rucio), VRE (k8s, Jupyter, REANA), OSSR: multi-science.
  • The hosted Rucio MCP already serves an escape site.
# escape.cfg — a real, shipped site config
[client]
rucio_host = https://vre-rucio.cern.ch
auth_host = https://vre-rucio-auth.cern.ch
oidc_issuer = escape
oidc_audience = rucio

Gateway + playbooks on the VRE: built once, serves LHC and astroparticle science.

Equitable inference

Who pays for inference?

  • Credit-card-per-student entrenches inequality, between groups and countries.
  • Host open-weight models on facility GPUs (NRP-style).
  • Bridge identity: WLCG token → scoped inference token (RFC 8693).

A facility responsibility: GPUs, model hosting, fair-share policy.

Agents already perform full HEP analyses end-to-end — Moreno et al., arXiv:2603.20179 — and that run kept stalling on usage rate-limits: access is already the binding constraint.

Deep-dive · cost

Per analysis: tens to hundreds of dollars

Illustrative: ~1M in + 1M out tokens per analysis.

Frontier (hosted)

$100–200+

Mid-tier (hosted)

$15–35

Open-weight (self-hosted)

GPU time only

Self-hosted open weights: ~GPU time. The equity lever.

Tiers from per-Mtok list prices (input + output) via the AI token pricing calculator, ballpark; the 1M + 1M budget is illustrative. Agents already run analyses end-to-end: Moreno et al., arXiv:2603.20179.

Research questions

What we don’t know yet

How should facilities expose context?

How do we share it?

Evaluating an autonomous operator?

What makes environments portable?

Who pays for inference?

What standardizes, across sciences?

Context, identity, trust, equity: the facility’s homework.

Thank you

Two trust problems, one facility layer

On: trust earned in git. Through: one brokered gateway, ~800 users, no client secrets.

Which parts do we build together?

Giordon Stark: AMG Weekly · PyHEP.dev 2026 · IRIS-HEP retreat · CLARIPHY

af-mcp-platform · rucio-mcp · ami-mcp · usatlas/marketplace · AF AI docs · ssh-oidc · pixi.sh · Gardner · Stark · Vukotic · Hu · UChicago · WLCG OTF #12 · 2026‑08‑25.
35ec3ac · 2026-08-24