~/posts/secrets-scanning-zero-dev-setup

Scanning every push in a 200+ repo GitHub org for secrets, with zero developer setup

14 min read 2730 words
cybersecuritycloudcorporatedevops

tl;dr

A GitHub push webhook triggers a Lambda that checks out the pushed commit, runs TruffleHog with verification, and stores only SHA-256 hashes of what it finds. Alerts go to Slack and a read-only dashboard. The baseline scan found about 100 secrets; the average fix took under three hours.

Our GitHub organisation at work has 200+ repositories and hundreds of pushes a day. With that many commits, secrets get committed. Not because people are reckless: people are busy, juniors are learning, and legacy repos have a surprising gravitational pull. A lot of our work is customer projects, so a leak isn’t just embarrassing. It’s contractual risk, breach risk and reputation risk.

So we built a scanner that runs on every push, on every branch, in every repo in the org, and asks nothing of developers. No local hooks, no CI changes, no “install this tool and keep it updated”, no merge blocking. It detects, tells the right person, and makes fixing it boring.

Boring is the goal. Nobody wants adrenaline in incident response.

Why “just add a tool” doesn’t work

With enough commits, someone will commit an API key, a cloud credential, an SSH key or a private certificate. That part is simple. The messy part is everything around it.

People ship under deadline pressure. Interns and juniors do what the codebase teaches them, and legacy repos are excellent teachers of bad habits. And security often arrives as “something we add later”: a tool rollout that needs buy-in from every developer, every repo owner and every CI pipeline.

That collapses at scale. If the scanner depends on developers opting in, the repos that need it most will be the last to adopt it.

What I wanted instead:

  • Automatic and org-wide, invisible to developers.
  • Fast feedback: minutes, not days.
  • Cheap, so nobody starts negotiating with security about whether scanning pushes is “worth it”.
  • Ours. Not in the geopolitical sense: our logic, our data model, our integrations. If it breaks, we fix it. If we need a tweak, we don’t wait for a vendor roadmap.

And what it would deliberately not do: block pushes or merges, scan anything outside the org’s GitHub repos, or ask anyone to install something locally or retrofit CI on 200+ repos.

Security by design, without security by paperwork.

What gets scanned, and what counts as a secret

In scope: every repository in the org, every push event, every branch. The unit of work is the repository checked out exactly at the pushed commit SHA. That gives a clear, reproducible job: “scan this snapshot”.

Out of scope: anything not pushed to the org’s repos. No developer laptops, no external Git hosting, no artifact registries, no build logs. That’s a separate project with a separate risk model.

A secret is an API key, an SSH key, a TLS certificate or private key, a cloud credential (AWS, GCP), a mailing service credential, or a high-entropy string that’s probably a password. Any secret in git history is unacceptable. There’s no safe window where it was “only in a branch”. Git history is a very effective long-term storage system, and that’s a compliment only until it isn’t.

The whole system in one line

GitHub push webhook → the receiver checks it’s really from GitHub → skip if this commit SHA was already scanned → an AWS Lambda clones the repo at that SHA → TruffleHog scans the files and verifies what it finds → findings are normalised → only metadata and SHA-256 hashes are stored → Slack alert and a read-only dashboard.

That’s all of it. It’s deliberately direct.

GitHub org push

Webhook receiver

Idempotency commit SHA

Lambda scan job

git clone depth=1

checkout commit SHA

TruffleHog filesystem scan + verify

JSON findings

Hash-only findings store

Slack alert

Read-only dashboard

The same flow in time order, from push to alert:

Internal DashboardSlackFindings StoreTruffleHogGit (clone/checkout)AWS Lambda (scan)Webhook ReceiverGitHubInternal DashboardSlackFindings StoreTruffleHogGit (clone/checkout)AWS Lambda (scan)Webhook ReceiverGitHubpush webhook (event + signature)verify HMAC signatureidempotency check(commit SHA)invoke scan(commit SHA, repo)clone --depth 1repo workspace at commitscan filesystem + verifyfindings JSONupsert current findings (hash-only)notify (no secret values)dashboard reads current state

Where it runs

The scan jobs run in AWS Lambda in eu-west-3 (Paris), with 1024 MB of memory each and a global concurrency cap. The cap matters because pushes are bursty: one busy hour can hold a large slice of the day’s pushes.

The scanner scales up or down without developers changing anything. That’s the entire point.

The price is time to detection. The earlier stateful proof of concept (more on it below) was faster. With Lambda, a push takes about 40 seconds to turn into an alert. In practice that’s fast enough to keep the human feedback loop tight, without making operations complicated.

The webhook receiver: is it real, and have we seen it?

The receiver has two jobs, and both must be boring and correct.

Is it really GitHub? GitHub signs each delivery with an HMAC of the payload, and we verify that signature. If it’s invalid, the request is rejected. We don’t rely on IP allowlists, a WAF or other perimeter controls. Those aren’t bad, but for this system the signature is the primary control, and it’s straightforward to check correctly.

Have we already scanned it? Webhooks can be delivered more than once, and retries are normal. So the pushed commit SHA is the dedupe key, which makes the handler idempotent: a duplicate delivery of an already-scanned SHA does nothing.

The logic, in intentionally boring pseudocode:

function handlePushWebhook(req) {
  if (!verifyGithubHmac(req)) return 401;

  const sha = req.payload.after; // pushed commit SHA
  if (alreadyScanned(sha)) return 200; // no-op

  markScanned(sha);
  invokeLambdaScan({
    sha,
    repo: req.payload.repository.full_name,
    branch: req.payload.ref,
  });

  return 202; // accepted
}

The only state kept here is which SHAs were scanned. It contains nothing sensitive.

The scan job

Each scan is stateless and repeatable:

  1. git clone --depth 1, to keep clone time and storage low.
  2. Check out the pushed commit SHA.
  3. Run TruffleHog in filesystem mode, with verification on.
  4. Read its JSON output, normalise the fields, write findings to storage.
  5. Send a Slack notification, metadata only.
  6. The dashboard reads the current open state.

The repo is scanned as a snapshot of files, not by walking historical git objects. The .git directory is excluded so object data isn’t scanned, and .gitignore is respected to cut noise from generated files.

That keeps the scan aligned with the question we care about: “did someone just push a secret that now exists in the repo at this commit?” Scanning full history is possible. It’s just a different problem, a different runtime profile, and often a different budget.

TruffleHog, unmodified

We run an unmodified TruffleHog binary, built from the latest release at build time when there’s a newer one. Verification is on for the detectors that support it: TruffleHog tries the credential against the provider to see if it’s live, which cuts false positives. The settings, in plain words:

  • Filesystem scanning: scan what the repo looks like at the commit.
  • JSON output: makes storage and dedupe manageable.
  • Exclude .git: don’t rescan blobs.
  • Respect .gitignore: less noise.
  • No custom detectors yet: when we find a gap, we prefer contributing upstream over maintaining a fork.

That last point surprises people. Owning your tooling doesn’t mean forking everything and carrying patches forever. Often it means owning the integration and the system around the tool, and fixing the tool upstream when it’s missing something.

Storing findings without storing secrets

The store only holds current open findings. When a finding is resolved, its row is deleted. That’s deliberate.

The schema:

CREATE TABLE IF NOT EXISTS secrets (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    repo TEXT NOT NULL,
    branch TEXT,
    file_path TEXT NOT NULL,
    provider TEXT,
    hash TEXT,
    first_seen_at TEXT NOT NULL,
    last_seen_at TEXT NOT NULL,
    resolved_at TEXT
)

What’s stored: the repo name, the branch from the push, the file path, the provider label (which detector flagged it), a SHA-256 hash, and first/last seen timestamps.

What’s never stored: the secret itself. Not in the database, not in Slack, not in the dashboard.

To compute the hash, the secret sits in memory for as short a time as possible: hash it, keep the hash, drop the value. A cryptographic hash gives a stable identifier for dedupe and for “is this the same secret coming back?”, without keeping a credential anyone could use.

A finding’s life:

first detection

still present (last_seen_at updated)

no longer detected

row dropped (current-state only)

Open

Resolved

Deleting resolved rows has a clear downside: without a separate event log, history is gone. We can’t answer “how long did that exact secret live?”. We accepted that, because the system’s job is fast remediation, not analytics. If we want analytics later, an append-only event log can be added next to the current-state store without changing it.

Alerts and the dashboard

Alerts go to Slack, and a read-only internal dashboard lists what’s currently open, filterable by repo and provider, so owners get a clear list of what to fix.

A Slack alert has enough context to act, and nothing that leaks the secret. A sanitised example:

🚨 Secret detected

Repo: customer-project-api
Branch: feature/onboarding
Commit: 1a2b3c4 (pushed by j.doe)
File: config/dev.env
Provider: AWS (verified)

Next steps: remove from repo + rotate/revoke if needed
Dashboard: internal read-only link

The “verified” marker matters. It short-circuits the “probably a false positive, I’ll ignore it” reflex, and that reflex is how small leaks become incidents.

Who gets pinged, and what they do

Routing is intentionally boring. The lead dev on the project is responsible, and the commit author is included for context, through an internal directory mapping plus the commit’s author. The goal isn’t to shame anyone. It’s to put the alert in front of the person who can fix it fastest.

Severity is boring too: everything is critical. We don’t play “maybe this secret is low value”. Once a secret is in git history, the safe assumption is that it’s compromised.

The fix is what you’d expect, and that’s the point:

  1. Remove the secret from the repository.
  2. Rotate or revoke the credential, depending on the provider and how likely it is to be compromised.
  3. Check the scanner no longer sees it: the finding disappears and its row is dropped.
  4. Stop it happening again by moving secrets into a secrets manager or environment injection.

There are no automated fix PRs today, on purpose. Automation can help, but it can also add friction and complexity. Detect and route came first, because that alone removes most of the risk quickly.

What happened after rollout

The first scan across the org, the baseline, found about 100 secrets. Most were fixed quickly. The slowest fix took one business day, and the average was under three hours.

The most important result was cultural. The system taught developers what “good” looks like by giving feedback right away. It’s much easier to learn from a Slack alert ten minutes after a mistake than from an incident report two months later.

It also confirmed something practical: the repos that leak are usually the oldest, the most copy-pasted, and the most “just get it working” in local dev. That’s not a moral failure. It’s a sign those repos need better patterns and templates.

Findings per week

Findings per week
loading chart…
Weekly new vs resolved findings.

Open findings at week end

Open findings at week end
loading chart…
Illustrates open findings trending down after baseline cleanup.

Time to remediate

Mean time to remediate
loading chart…
Mean time to remediate distribution (buckets).

Top leaked providers

Top leaked providers
loading chart…
Top provider categories from findings.

Repo hygiene

Repo hygiene
loading chart…
Clean repos vs repos with open findings.
Clean repo rate
loading chart…
Share of repos with no open findings.

The first version: fast, with a slow-motion footgun

Before this, we built a proof of concept on a small EC2 instance (t4g.micro) with a stateful workspace:

  • a webhook receiver,
  • a simple FIFO queue backed by a database table,
  • a worker process,
  • a per-repo lock,
  • a cached workspace per repo, updated with git pull,
  • SQLite for state,
  • Slack notifications for what changed.

It was fast: about 5.5 seconds per scan on average, and 12.76 seconds at the 95th percentile. The persistent workspace also helped us find edge cases and contribute fixes upstream.

PoC average: 5.5 sPoC average5.5 sPoC p95: 12.76 sPoC p9512.76 sLambda: 40 sLambda, push→alert~40 s
The proof of concept’s numbers are scan times; the Lambda number is the whole trip from push to alert. Roughly 7× slower on average, and still well within “minutes, not days”.

The problem was operations. The cache directory needed cleanup. Cleanup got missed. Storage grew beyond expectations. The system was still “working”, but the operational risk wasn’t acceptable. That’s the classic problem with stateful workers: performance is great until the state starts having opinions about your disk.

So we moved to stateless scanning on Lambda. It’s slower, about 40 seconds from push to detection, but it’s simpler to run and scales cleanly with load. In a security system, predictable and boring is often worth more than fast and clever.

The scanner’s own security

This system isn’t a perimeter fortress. It’s a targeted control with a specific threat model:

  • Webhook authenticity comes from GitHub’s HMAC signature.
  • Replays and duplicate deliveries are absorbed by the commit-SHA dedupe.
  • The GitHub token used for cloning is a fine-grained personal access token with read-only permissions.
  • There’s no IP allowlist or WAF in front of the receiver. HMAC is the primary control.

That last one is a trade-off. For defence in depth, perimeter controls can be added later; nothing in the design prevents it. We just didn’t make them a prerequisite for shipping a useful scanner.

What it costs

Lambda is billed by memory × time, plus a small fee per request. At 1 GB and about 40 seconds per scan:

compute per push:  1 GB × 40 s = 40 GB-seconds
cost per push:     40 × (price per GB-second) + request cost

In practice it’s close to zero for us, because it fits in the AWS free tier.

What might come next: automatic fix PRs

The obvious next step is for the system to open a PR that removes the secret and replaces it with an environment reference, possibly with help from LLM agents. That could be valuable, especially in a large org where the same mistake repeats across repos.

It stays on the roadmap for a reason. Auto-remediation easily becomes a source of friction if it’s noisy or intrusive, and the current system already delivers most of the value by detecting, routing and keeping the feedback loop tight. If we add it, it should be opt-in, conservative, and approved by a human.

hash: c2e
EOF