How do you test a Content Security Policy in CI and staging?

A three-layer setup for catching CSP regressions before users do: the policy diff in code review, a scored scan that fails the build on a preview URL, and report-only on staging and production for the traffic your pipeline will never generate.

Published

Most CSP rollouts have a gap between “it worked when I browsed the site” and “it is enforced in production”, and the gap is where the regressions live. Someone adds a chat widget, the policy gains a wildcard to make it work, and nobody notices until an audit six months later points out that script-src now allows half the internet.

The setup that actually holds has three layers, and they catch different things. The policy diff in code review catches deliberate weakening. A scored scan in CI, run against a preview or staging URL, fails the build on typos and unsafe directives before anyone merges. Report-only against real traffic catches the pages and browsers your pipeline will never visit. Skip any one of the three and you will find out which one you skipped.

None of this needs a browser automation rig. Two of the three layers are a curl call and a header.

Layer one: treat the policy as code

If your CSP is a string in a config file, a reverse proxy config or a middleware, it is already reviewable. Most teams never review it, because the diff looks like line noise and the reviewer has no way to judge whether https://*.vendor-cdn.net is reasonable.

Make it reviewable by splitting the policy into named groups in source, one per reason, with a comment saying who asked for it and when:

const scriptSrc = [
  "'self'",
  "'nonce-{NONCE}'",
  "'strict-dynamic'",
  // Payments vendor, added 2026-03 by the checkout team. Removing this breaks
  // the card form. Reviewed with them 2026-09, still required.
  'https://js.payments-vendor.example',
];

Then the diff says something. A new entry with no comment is a conversation in review rather than an archaeology project next year. This costs nothing and catches the single most common way policies rot, which is not a bug but a deliberate, reasonable-at-the-time loosening that never got revisited.

One rule worth enforcing by convention: 'unsafe-inline' and 'unsafe-eval' do not enter script-src without a named owner and a removal date. Otherwise they are permanent.

Layer two: score the built output in CI

This is the layer people ask about and the one most pipelines are missing.

The important thing is what you point the check at. Not your dev server. A dev server injects inline scripts for hot module reload, which is why developers add 'unsafe-inline' in the first place, and it usually serves different bundle filenames than the production build. Point the check at a production build: a preview deployment, a staging URL, or a container started from the real artifact.

What a scan catches that a browser does not:

  • Typos. A script-scr directive is not an error. It is an unknown directive, and browsers ignore unknown directives silently. Your site works perfectly and your policy has a hole exactly the size of that directive. The browser will never tell you.
  • A policy permissive enough to never break anything. default-src * passes every functional test you can write. That is the problem with functional tests as your only CSP gate.
  • JSONP and known bypass endpoints on origins you have allowlisted, where an allowed host serves an endpoint that will execute attacker-supplied callbacks.
  • The header missing entirely, because someone changed the proxy config.

For a manual check, the free CSP scanner takes a URL, fetches the live policy including Report-Only, and returns two 0 to 100 scores (security and quality) with severity-ranked findings and a fix for each. Above 80 is solid, 50 to 79 needs attention, below 50 means real gaps. No account needed, which makes it the fastest way to see whether this layer would find anything on your site before you build it into a pipeline.

To run it on every build you want the REST API, which is available from the Business plan (€129.99/month). There is no official GitHub Action to install. It is a curl call and a status code, which we prefer, because an action is one more thing to keep updated:

- name: CSP gate
  env:
    CENTRALCSP_TOKEN: ${{ secrets.CENTRALCSP_TOKEN }}
  run: |
    # Base URL, paths and field names: see the OpenAPI reference, not this snippet.
    SCORE=$(curl -sf -X POST "$CENTRALCSP_API/scans" \
      -H "Authorization: Bearer $CENTRALCSP_TOKEN" \
      -H 'Content-Type: application/json' \
      -d "{\"url\":\"$PREVIEW_URL\"}" | jq -r '.security_score')
    echo "CSP security score: $SCORE"
    [ "${SCORE%.*}" -ge 70 ] || { echo "CSP gate failed"; exit 1; }

That snippet is the shape of the thing, not a copy-paste: take the base URL, the scan path and the score field from the OpenAPI reference, because a response schema outlives any article quoting it. Workspace API keys act as the person who created them, with the same website roles and the same plan limits, and revocation is immediate, so a CI key is worth creating as a dedicated read-oriented user rather than under a founder’s account.

Two notes on thresholds. Pin a floor, not a target: a build that must score 90 gets disabled the first time a legitimate vendor is added. And fail hard on the categorical findings (missing header, unknown directive, 'unsafe-eval' appearing where it was absent) while only warning on score movement. A gate that fires on noise gets bypassed, and a bypassed gate is worse than no gate because everyone believes it is running.

If your pipeline is already talking to an assistant, the same data is available through the built-in MCP server on the same plan, so “has the score dropped since last week and which finding is new” is a question you can ask rather than a script you maintain.

Layer three: report-only, on traffic you did not generate

CI runs the pages your tests visit. Real users visit the others.

Ship the candidate policy as a Content-Security-Policy-Report-Only header with a reporting endpoint attached, and the browser reports what it would have blocked without blocking anything. It cannot take down checkout. That is the whole reason it exists, and it is the only layer that exercises the Safari-on-an-old-iPad path, the locale that loads a different consent vendor, and the admin page nobody wrote a test for.

Reporting-Endpoints: csp="https://MyEndpoint.report.centralcsp.com"
Content-Security-Policy-Report-Only: default-src 'none'; script-src 'self' 'nonce-{NONCE}' 'strict-dynamic'; report-uri https://MyEndpoint.report.centralcsp.com; report-to csp

On staging this gets you a fast signal from your own QA. On production it gets you the real distribution, which is the one that matters. We have written separately on how long to run report-only and the move from report-only to enforced, and the short answer is that you stop when the unexplained violations reach zero, not when a calendar says so.

For preview environments, the wrinkle is that the hostname changes on every branch. CentralCSP restricts which origins may post reports for a site through ingestion filters, which is what stops a stranger filling your dashboard with junk. Preview domains need that wildcard added deliberately, and it is worth pointing previews at a separate site entry rather than mixing ephemeral branch noise into the production site’s data and its report quota.

What each layer is actually for

LayerCatchesMisses
Policy diff in reviewDeliberate weakening, undocumented sourcesAnything nobody reviews carefully
Scored scan in CITypos, unsafe directives, missing header, bypass-prone originsBreakage on pages the scan does not load
Report-only on real trafficReal breakage, forgotten pages, browser differencesNothing, given enough time, which is the cost

The failure mode of layer two alone is a policy that scores 95 and breaks the checkout on Safari. The failure mode of layer three alone is a policy that never breaks anything because it permits everything. They are complementary, and neither replaces reading the diff.

Getting local development out of the way

Before any of this is worth automating, developers need to be able to iterate on a policy without a deploy, because a feedback loop measured in CI minutes produces 'unsafe-inline' by lunchtime.

The free Chrome extension does that locally, with no account: its Rewrite mode swaps a candidate policy in for the site’s real one, enforce or report-only, for your session only, so you can log in and walk through checkout under a policy that has not shipped. Build mode assembles a draft policy from a strict baseline while you browse. We covered that workflow in testing a CSP without deploying it.

Get that loop working first, then automate the gate. In the other order, the gate just becomes the thing people learn to route around.

Frequently asked questions

Can a CI pipeline catch every CSP problem?

No, and treating it as if it could is the usual mistake. CI runs the pages your test suite visits, in one browser, with no real users. It reliably catches a weakened policy, a typo in a directive name and a missing header, which is most regressions. It cannot catch the locale-specific widget on a page nobody wrote a test for. Real traffic in report-only covers that, and the two layers are not interchangeable.

Why does my CSP work locally but break in production?

Almost always because the dev build differs from the production build. Dev servers inject inline scripts for hot reload, so developers add unsafe-inline to get through the day and it never comes back out. Bundlers also emit different filenames and sometimes different CDN origins in production. Test the production build, not the dev server, and diff the policy your production pipeline actually emits.

How do I test a CSP on ephemeral preview environments?

Point the preview deployment at a report-only policy and a reporting endpoint, then let the branch collect for as long as it lives. The part people forget is ingestion filtering: preview domains change on every branch, so allow the wildcard origin for previews deliberately rather than letting any host post reports into your production site data.

Should the build fail on a CSP finding, or only warn?

Fail on things that are unambiguous and cheap to fix: a missing header, an unknown directive name, unsafe-inline or unsafe-eval appearing in script-src where they were not there before. Warn on score movements, because a legitimate new vendor can move a score a few points and a pipeline that cries wolf gets bypassed within a month.