Skip to content

Fixing an API Key Leaked Into Serverless Function Logs

10 min read · updated August 11, 2026

A key in a log line is worse than a key in a repository, because logs are shipped to more places than code is, retained by default, and read by tooling with wide access. The good news is that the mechanism is almost always one of four things, and none of them is somebody typing console.log(apiKey).

Revoke before you investigate

Do this before reading any further. The key is compromised; how it got there does not change that, and every minute of investigation is a minute the value is still valid. Revoke it in the provider’s console or management API, issue a replacement, and deploy the replacement. Only then work out what printed it.

Two things to check while revoking. Look at the provider’s usage dashboard for the period since the first appearance in logs — unexplained usage is the difference between an incident and a near-miss. And check whether the same key was used anywhere else; a key shared across three services means three deployments, and a partial rotation leaves one service broken and one key live.

Deleting the log group does not undo the exposure. Logs are commonly forwarded to a SIEM, an observability vendor, an S3 export or a third-party log platform within seconds of being written, and those copies have their own retention. Treat deletion as tidying up, never as remediation.

What actually printed it

Almost nobody logs a secret on purpose. These four print it as a side effect of doing something else, and they are worth recognising on sight.

  • An HTTP client error serialised whole. This is the most common by a wide margin. Popular HTTP clients attach the request configuration — including headers — to the error they throw, so logger.error(err) on a failed model call prints Authorization: Bearer sk-... along with the status code. The log line looks like an ordinary error report and contains a credential.
  • A deliberate environment dump. console.log(process.env) or a Python logging.debug(os.environ), added during a cold-start debugging session and never removed. Grep for it; it is a one-line fix and it recurs.
  • A framework startup banner. Some frameworks and configuration libraries print their resolved configuration at boot, which is exactly the merged view of your environment. It appears once per cold start, which makes it easy to miss and frequent enough to matter.
  • An unhandled exception carrying locals. Traceback formatters that include local variables — and several error reporting SDKs do this by default — will include the key if it was a local in the failing frame.

Note what these have in common: in every case the key was in memory legitimately, and the logging layer had no way to know one string was different from another. That is the reason redaction has to be structural rather than a matter of care.

Finding every occurrence

Search by the shape of the credential rather than by the value, because the value you found may not be the only one that leaked. Most provider keys have a recognisable prefix, which makes a regex practical. In CloudWatch Logs Insights, across every function’s log group:

fields @timestamp, @log, @logStream, @message
| filter @message like /(sk-[A-Za-z0-9_\-]{20,}|Bearer [A-Za-z0-9_\-]{20,})/
| sort @timestamp asc
| limit 200

Run it over the full retention period, not the last hour. The first timestamp is the one that matters: it bounds the exposure window and it usually lands on a deploy, which identifies the change that introduced it. Then find the code:

# The deliberate dumps.
rg -n "process\.env\b(?!\.)" --pcre2 src/
rg -n "os\.environ\b(?!\[)" --pcre2 src/

# The error-object logs, which are the ones that matter.
rg -n "logger\.(error|warn)\((err|error|e)\)" src/
rg -n "JSON\.stringify\((err|error)" src/

If the search is clean and the leak is real, look at your dependencies rather than your code — an error-reporting SDK with request-body capture enabled, or a middleware that logs request and response headers for outbound calls.

Who could read it

Establishing the audience determines whether this is an internal hygiene fix or a disclosure. Enumerate honestly:

  • Everyone with log read access. On AWS that is anyone holding logs:FilterLogEvents or logs:GetLogEvents on the log group, which in many accounts is a broad read-only role rather than a narrow one.
  • Every downstream sink. Subscription filters, Firehose deliveries, S3 exports and third-party agents. Each has its own retention and its own access list, and a key deleted from CloudWatch lives on in all of them.
  • Anything that indexed it. A log search platform may have the string in an index that survives deletion of the source document, and support tickets or Slack channels where somebody pasted a stack trace are a distribution channel too.

Check the log group’s retention while you are there: a group created without an explicit retention setting keeps its events indefinitely, and an indefinite retention on a log that once contained a credential is a permanent liability.

Making it structurally hard

  1. Never log an error object whole. Log the fields you want — status, message, request id, elapsed time — and never the config or the headers. This single rule removes the most common cause.
  2. Redact in the logger, not at the call site. A serialiser hook that masks any value matching a credential pattern, and any key named authorization, api_key, token or secret, catches the cases nobody anticipated. Every structured logging library supports this; it is usually one configuration block.
  3. Set explicit retention on every log group. A short retention limits how long any future mistake is readable, and it is a one-line IaC change.
  4. Detect rather than hope. A metric filter on the same regex you searched with, alarming on any match, turns the next occurrence into a page in minutes instead of a discovery in months.
aws logs put-metric-filter \
  --log-group-name /aws/lambda/inference-worker \
  --filter-name credential-shape \
  --filter-pattern '%(sk-[A-Za-z0-9_-]{20,})%' \
  --metric-transformations \
      metricName=CredentialInLogs,metricNamespace=Security,metricValue=1

The related mistake worth closing at the same time is the one on a secret baked into a Docker image, which has the same root cause — a credential present somewhere it can be copied — and a much longer half-life.