← Back to leaderboardProject evidence record

garak

This free page contains the context needed to interpret the published static score. Deep review and runtime validation are separate evidence levels, not hidden score modifiers or certifications.

78Static score
out of 100

Current posture summary

The pinned garak snapshot scored 78 under the static AI baseline. All four required evidence providers completed with no unexpected scope gaps; the result describes repository-visible posture and does not test runtime behavior.

Methodology: 2026-04-11.standards-static.v2

Repository: https://github.com/NVIDIA/garak

Dimension evidence

  • Repo posture31
  • Agentic guardrails100
  • AI data exposure100
  • Observability & auditability100
  • Evidence readiness82

Strongest and weakest areas

Strongest: Agentic guardrails (100)

Weakest: Repo posture (31)

Evidence-backed strengths

  • All four required static evidence providers completed successfully.
  • No unexpected static controls were left unassessed, and repository evidence supported the assessed agentic, data-exposure, observability, and evidence-readiness controls.

Evidence-backed weaknesses

  • Multiple workflows reference third-party GitHub Actions by mutable version tags rather than immutable full commit SHAs.
  • The repository-posture dimension scored 31, lower than its other static dimensions in this snapshot.

Limits and next action

  • Twelve expected controls outside static reach—including prompt injection, authorization, isolation, exfiltration, and resource-limit behavior—were not assessed.
  • Dependency-update automation and workflow-permission observations were not promoted into public findings because their evidence or severity was not sufficiently established for publication.
  • Documentation, configuration, and code markers are not proof that controls remain effective in a deployed runtime.
  • Editorial adjudication was AI-assisted project-owner review, not independent certification or human ground truth.

Remediation guidance

  • Pin third-party GitHub Actions to reviewed full commit SHAs while retaining readable release tags in comments.
  • Validate dependency-update coverage and least-privilege workflow permissions with exact workflow evidence before publishing further claims.
  • Use selective runtime testing to evaluate prompt-injection resistance, tool authorization, secret isolation, exfiltration boundaries, output handling, resource limits, and security telemetry.

No reviewed public case file is attached.

Public-safe finding summaries

medium

Build workflow dependencies use mutable references

Repository workflows reference third-party Actions by mutable version tags, including actions/checkout@v3 and actions/setup-python@v4, rather than immutable full commit SHAs. This is a static build-integrity weakness, not evidence of compromise.

.github/workflows/docs.yml:29.github/workflows/docs.yml:31.github/workflows/release.yml:28
Provenance

Source run: run_garak_538f0407-d3c7-4d1f-9c4f-b62a2aefb0ee

Source schema: executive_summary.v1

SHA-256: 14544e5fa0f90dda7bb80969cdae89ee5ec4aac93235a7e0352a1dc41aa6655d

Independent commercial boundary

Need a fresh, private, or deeper assessment?

thethermark can expand scope, evidence, runtime validation, monitoring, and reporting. Payment never predetermines or changes the public result.

Request deeper assessment