Selected work

Teaching an agent that severity isn’t difficulty

An agentic LLM framework for vulnerability remediation in NASA’s Deep Space Network, built during my summer internship at the Jet Propulsion Laboratory.

Role
AI & Cybersecurity Intern, NASA/JPL, Jun – Aug 2026
Mentor
Mark D. Johnston, JPL/Caltech
Built with
Python, Pydantic AI, MCP, FastAPI, React/TypeScript
Report
Technical report (PDF)

The problem

The Deep Space Network is NASA’s array of large radio antennas in California, Spain, and Australia, and the link to every spacecraft operating beyond the Moon.1 A missed communication pass can mean lost science data, so the network’s most important property is availability.

That makes routine security work delicate. The standard fix for a vulnerability is to patch and reboot, and rebooting a mission-critical machine during an active tracking pass is not an option. Scanners rank findings by CVSS,3 a severity score that is deliberately context-free: it can’t know what a machine does, or whether it can go down tonight. Analysts fill that gap by hand, cross-checking every alert against documentation, configuration records, and years of institutional knowledge.

Meanwhile the number of disclosed vulnerabilities keeps climbing, and LLMs are beginning to speed up how attackers find and exploit them. Defense that depends entirely on manual correlation can’t keep pace, and correlation is exactly what a well-equipped LLM agent is good at.

What I built

An end-to-end pipeline that turns a raw scanner finding into a reviewed, structured remediation recommendation. The agent decides which tools to call to gather context, reasons about the vulnerability, and returns a validated assessment. It recommends; it never acts. Step through it below, and switch between the two example alerts.

Example alert · synthetic data

Scanner severity Remediation difficulty A · Higheasy B · Medhard

Alert

A vulnerability scanner reports a finding with a CVSS score, a universal severity number. It knows nothing about what the machine does for the Deep Space Network, or what applying the fix would interrupt.

Outdated web-server package
AlertSYN-A
Hoststatus-web-02
Severity8.8 · High
Kernel privilege-escalation patch
AlertSYN-B
Hostsp-host-14
Severity5.5 · Medium

Context

Instead of stuffing every document into the prompt, the agent calls MCP tools and fetches only what it needs: engineering documentation first, then the wiki, the acronym database, and the configuration database.

  • search_docs("status-web-02 role")Public status page. Redundant pair behind a load balancer.
  • config_item("status-web-02")Non-mission system, standard maintenance window.
  • wiki_search("package update rollback")Rolling update, one node at a time. Nothing used for tracking restarts.
  • search_docs("sp-host-14 role")Processes downlink signal during active tracking passes.
  • config_item("sp-host-14")Mission-critical. No hot spare.
  • wiki_search("kernel patch procedure")Requires a reboot. Schedule outside tracking passes.

Agent

A Pydantic AI agent has to fill a typed schema, not write prose. When its output breaks the schema, the validation error goes back to the model and it tries again, at temperature zero.

  1. Attempt 1Valid on the first try.
  1. Attempt 1difficulty: "medium" is not one of "easy" or "hard". Error sent back to the model.
  2. Attempt 2Valid. Passed to output.

Output

The result is validated JSON: a difficulty class based on operational feasibility, the reasoning behind it, deployment considerations, and the documents it relied on. Whether a flaw is known to be exploited comes from CISA's KEV catalog2 as a hard fact, not a model guess.

{
  "vuln_id": "SYN-A",
  "severity": "High",
  "difficulty": "easy",
  "reason": "Routine package update on a
    redundant, non-mission host. Rolls back
    cleanly; fits a normal maintenance window.",
  "considerations": ["Update one node at a time"],
  "known_exploited": false
}
{
  "vuln_id": "SYN-B",
  "severity": "Medium",
  "difficulty": "hard",
  "reason": "Needs a reboot of a mission-critical
    host with no hot spare. Schedule outside
    tracking passes.",
  "considerations": ["Coordinate with the tracking
    schedule", "Verify signal processing after reboot"],
  "known_exploited": false
}

Review

Nothing is applied automatically. An analyst reviews each recommendation in a dashboard, grouped by difficulty, and decides. The system recommends; people act.

AlertScanner severityAgent difficulty
SYN-AHigheasy
SYN-BMediumhard

Sorted by severity, A comes first. Sorted by what the fix costs the network, B is the one that needs planning.

Stage 1 of 5 · Alert
Figure 1. The vulnerability-remediation pipeline I built at NASA/JPL, shown with two made-up alerts. Scanner severity and remediation difficulty point in opposite directions: the agent ranks by what a fix would cost the network, and a person reviews every recommendation. The alerts, host names, tool names, and tool results are illustrative. The real system ran on internal DSN data, which isn’t shown here.

The central result is that inversion. A high-severity flaw whose fix is a routine, easily rolled-back update is classified easy. A moderate one whose fix needs a reboot of core services, a firmware change, or a performance-costing patch is classified hard. The agent gets there by consulting the same internal documentation an analyst would, and it has to cite what it used.

Structured output as a contract

Free-form prose can’t feed a pipeline. The agent is built on Pydantic AI,4 which constrains the model to fill a strictly typed schema. When the output violates it, the validation error is fed back and the model corrects itself, up to a fixed number of retries. That loop is what turned a promising demo into a component other tools can rely on. Inference runs at temperature zero, and every assessment is stamped with the prompt and schema versions, the model ID, and the backend it ran on, so any result can be traced and reproduced.

Model temperature 0 Validate against the typed schema Assessment valid JSON pass fail: the validation error is fed back, then retry
FieldWhat it holds
severityStandard severity level, from the scanner
difficultyOperational feasibility: easy or hard
reasonWritten rationale for the classification
impact_analysisAffected systems and mission impact
considerationsDeployment and remediation considerations
remediation_action_commandThe exact fix, or null if it’s manual
verificationGenerated tests to confirm the patch and service health
known_exploited, kev_*CISA KEV enrichment, injected as fact
Figure 2. The self-correcting loop, and the principal fields of the remediation assessment from the report’s appendix.

Context on demand, through MCP

Putting every potentially relevant document into the prompt would be expensive and would overflow the context window. Instead, each context source sits behind a Model Context Protocol5 server, and the agent retrieves only what a given vulnerability needs. The sources are the engineering documentation (the primary, authoritative source), an internal wiki of procedures and known issues, an acronym database, a configuration-item database, and the existing vulnerability data. Each server runs as a local process that talks only to internal resources, so sensitive data stays inside approved boundaries.

The backend is model-agnostic. Switching between three approved backends takes one configuration value, so analysis can be routed by data sensitivity, and the tool improves as the underlying models do.

Context sourcesvia Model Context ProtocolAgentStructured outputstyped, cited, auditableReviewEngineering docssystem referenceInternal wikiprocedures, known issuesAcronym databasesystem-name lookupConfig-item databasedeployed assetsVulnerability reportsinternal scanningCISA KEVactively exploitedAI agentpicks its own contextPydantic AIschema-validated output3 swappable backendsClassificationwith mission impactProposed fixJSON + commandKill chainattack-path diagramHuman inthe loopauthorizes everyrecommendationCaching, checkpoints, adaptive pacing, and cost tracking wrap every model call.

1 · Context sources

Engineering docsInternal wikiAcronym databaseConfig-item databaseVulnerability reportsCISA KEV
retrieved on demand

2 · Context connectors

MCP serverslocal processes that only talk to internal resources
only the context this vulnerability needs

3 · Reasoning

Pydantic AI agenttyped schema, self-correcting
Three swappable backendsswappable in config
the agent returns validated, cited JSON

4 · Outputs

Classificationwith mission impact
Proposed fixJSON + command
Kill chainattack-path diagram
every recommendation

5 · Human in the loop

Dashboardan analyst reviews and decides; nothing is applied automatically

Caching, checkpoints, adaptive pacing, and cost tracking wrap every model call.

Figure 3. System architecture: a Python analysis engine, a FastAPI service that runs and streams analyses, a React/TypeScript dashboard, and context connectors implemented as MCP servers.

Chains, not lists

Attackers rarely use one vulnerability. They chain them,6 turning a minor foothold into a position from which a serious compromise becomes possible, and a flat list of findings hides that. A second agent reasons about how vulnerabilities across interdependent systems could combine, and the dashboard draws the result as an interactive attack-path diagram.

Attacker’s patheach step builds on the last
  1. Low severity
    Web foothold
    A minor flaw on an exposed service gets an attacker in.
  2. Medium severity
    Privilege escalation
    A second, unrelated flaw raises their access on that host.
  3. Enabled by the chain
    Lateral movement
    Dependencies between systems carry them further in.
  4. Critical
    Critical compromise
    Even though none of the individual findings looked urgent.
Figure 4. An example exploit chain, generalized from the report. Chain length and the analysis’s tool budget are configurable, and the capability is experimental.

Built to run for real

Caching
Repeated context queries are cached, cutting cost on repeated queries by an estimated 50–70%.
Checkpoints
Progress is saved after every assessment, so an interrupted run resumes and loses at most one item.
Adaptive pacing
When a backend rate-limits, the delay between requests widens and honors cooldown hints, then relaxes as pressure clears.
Cost tracking
Input, output, and cached tokens are recorded per assessment, with actual and cloud-equivalent cost.

What comes next

The next step is measuring how good the recommendations really are, by comparing the agent’s calls with what experienced DSN analysts would decide. Doing that well takes a purpose-built dataset: real vulnerabilities from the network, each paired with an analyst’s judgment and the reasoning behind it. Assembling that takes real analyst time and care, which is why it comes after the prototype rather than alongside it. Until it exists, I’m not putting an accuracy number on the system.

Two capabilities are already prototyped and ready for the same hardening: vision, so the agent can read the diagrams in DSN documentation directly, and a tightly sandboxed container that tries to reproduce an exploit to see whether it’s actually feasible.

Two research questions came out of the internship: whether structured vulnerability-test data improves the agent’s grasp of real-world impact, and whether it can reason over firewall rule sets to prioritize by actual reachability7 instead of assuming worst-case connectivity.

Thanks to my mentor Mark D. Johnston, deputy mentors Jeff Singer and Wesley Walker, all of JPL/Caltech, and to DSN Operations Group 4023. This work was done at JPL under a contract with NASA and supported by the JPL Summer Internship Program.
Pennel, B. (2026). Automating Cybersecurity Recommendations for Remediation of Vulnerabilities within NASA’s Deep Space Network: An Agentic Approach. Technical report, NASA Jet Propulsion Laboratory, California Institute of Technology. Full report (PDF, 11 pages).

References

  1. NASA Jet Propulsion Laboratory. Deep Space Network.
  2. Cybersecurity and Infrastructure Security Agency. Known Exploited Vulnerabilities Catalog.
  3. Forum of Incident Response and Security Teams. Common Vulnerability Scoring System (CVSS).
  4. Pydantic. Pydantic AI documentation.
  5. Anthropic. Model Context Protocol.
  6. Hutchins, E. M., Cloppert, M. J., & Amin, R. M. (2011). Intelligence-Driven Computer Network Defense Informed by Analysis of Adversary Campaigns and Intrusion Kill Chains. Lockheed Martin.
  7. The MITRE Corporation. MITRE ATT&CK.