← Back to leaderboardProject evidence record

smolagents

This free page contains the context needed to interpret the published static score. Deep review and runtime validation are separate evidence levels, not hidden score modifiers or certifications.

81Static score
out of 100

Current posture summary

The pinned smolagents snapshot scored 81 under the static AI baseline. All four required evidence providers completed with no unexpected scope gaps; the result describes repository-visible posture and does not test runtime behavior.

Methodology: 2026-04-11.standards-static.v2

Repository: https://github.com/huggingface/smolagents

Dimension evidence

  • Repo posture44
  • Agentic guardrails100
  • AI data exposure100
  • Observability & auditability100
  • Evidence readiness82

Strongest and weakest areas

Strongest: Agentic guardrails (100)

Weakest: Repo posture (44)

Evidence-backed strengths

  • All four required static evidence providers completed successfully.
  • No unexpected static controls were left unassessed, and repository evidence supported the assessed agentic, data-exposure, observability, and evidence-readiness controls.

Evidence-backed weaknesses

  • Two references in the test workflow use mutable GitHub Actions version tags rather than immutable full commit SHAs.
  • The repository-posture dimension scored 44, lower than its other static dimensions in this snapshot.

Limits and next action

  • Twelve expected controls outside static reach—including prompt injection, authorization, isolation, exfiltration, and resource-limit behavior—were not assessed.
  • A generated lockfile finding was suppressed during editorial adjudication because a reusable Python library is not necessarily expected to commit an application lockfile.
  • Documentation, configuration, and code markers are not proof that controls remain effective in a deployed runtime.
  • Editorial adjudication was AI-assisted project-owner review, not independent certification or human ground truth.

Remediation guidance

  • Pin third-party GitHub Actions to reviewed full commit SHAs while retaining readable release tags in comments.
  • Use selective runtime testing to evaluate prompt-injection resistance, tool authorization, secret isolation, exfiltration boundaries, output handling, resource limits, and security telemetry.

No reviewed public case file is attached.

Public-safe finding summaries

medium

Build workflow dependencies use mutable references

The test workflow references actions/checkout@v6.0.3 and actions/setup-python@v6 rather than immutable full commit SHAs. This is a static build-integrity weakness, not evidence of compromise.

.github/workflows/tests.yml:24.github/workflows/tests.yml:26
Provenance

Source run: run_smolagents_b3cc01e5-c526-4b6c-8490-8f80f5f010af

Source schema: executive_summary.v1

SHA-256: 4f09567160569700a406ea0997bff9b24303ed7321e5ea1f7358c655ac4d73a7

Independent commercial boundary

Need a fresh, private, or deeper assessment?

thethermark can expand scope, evidence, runtime validation, monitoring, and reporting. Payment never predetermines or changes the public result.

Request deeper assessment