Skip to content

Secret protection (Auth module) ​

Protection at rest of the secrets GrydAuth must be able to read back. Configuration section: GrydAuth:Security:SecretProtection Design decisions: ADR 0005

1. What this protects, and what it does not ​

SecretColumnWhy it must be reversible
User TOTP seedMfaFactors.EncryptedSecretThe seed is needed to compute and verify the code
Enterprise IdP client secretIdentityProviderConnections.ClientSecretEncryptedIt is sent to the IdP on every token exchange

Passwords are not here and never will be. They are hashed one-way with Argon2id (GeraltArgon2idPasswordHasher) and are never recovered — login compares hash against hash. These two are a different category: the plaintext genuinely has to come back, so they are encrypted, not hashed. Do not "harden" this by hashing them; it would end TOTP and enterprise SSO outright.

2. Configuration ​

jsonc
"GrydAuth": {
  "Security": {
    "SecretProtection": {
      "Scheme": "KeyRing",                      // KeyRing (default) | KmsEnvelope
      "ActiveKeyId": "gryd-secret-2026-07",
      "Keys": {
        "gryd-secret-2026-07": "${GRYD_SECRET_PROTECTION_KEY}"
      }
    }
  }
}

Generate a key:

bash
openssl rand 32 | basenc --base64url | tr -d '='
PropertyRequiredNotes
SchemenoKeyRing (keys from configuration) or KmsEnvelope (per-record data key wrapped by a KMS-held KEK).
ActiveKeyIdyes (KeyRing)Id of the key new secrets are protected with. Must be present in Keys.
Keysyes (KeyRing)keyId → 32-byte Base64URL key. Retired ids stay here until the re-wrap finishes.
Kms:Provideryes (KmsEnvelope)AzureKeyVault.
Kms:KeyIdyes (KmsEnvelope)Full key identifier, e.g. https://vault.vault.azure.net/keys/gryd/<version>.
Kms:DataKeyCacheLifetimenoDefault 00:10:00, hard ceiling 00:15:00.
Kms:DataKeyCacheSizenoDefault 1024 unwrapped data keys held in process.

Environment-variable form: GrydAuth__Security__SecretProtection__ActiveKeyId, GrydAuth__Security__SecretProtection__Keys__<keyId>.

This is validated at startup, in every environment ​

There is no Development exemption. A host without a usable key does not boot, and the failure message names the section and shows the command that generates a key.

That is deliberate. The previous validator let Development start with an empty section, so the first write failed instead — as an HTTP 409 Conflict whose message mentioned MFA, returned from an endpoint that was creating an identity provider. Configuration that is mandatory has to fail where it is cheap to diagnose.

3. How a protected value is built ​

g2.<scheme>.<keyRef>.<nonce>.<tag>.<ciphertext>
  • g2 — magic + format version.
  • scheme — kr (configured key ring) or kms (KMS envelope).
  • keyRef — opaque to everything except the owning scheme: for kr it is the key id; for kms it is base64url(kekId)~base64url(wrappedDataKey).
  • the last three segments are unpadded Base64URL.

Everything before the ciphertext is public metadata. It is authenticated but not encrypted, so altering any of it fails the AEAD tag.

The associated data is the whole point ​

g2 | scheme | keyRef | purpose | context

purpose and context come from the SecretProtectionScope the caller passes:

Purpose tokenContextBound to
mfa.totp-secretuser:<userId>/factor:<factorId>the user and the factor
federation.idp-client-secrettenant:<tenantId>/connection:<connectionId>the tenant and the connection row

The associated data is never stored — it is rebuilt on read from the row's own identity. So a ciphertext only opens in the exact place it was written:

  • a TOTP seed moved into the client-secret column fails;
  • one tenant's client secret placed in another tenant's connection fails;
  • a seed lifted onto another user's factor fails.

Before this design, all three silently succeeded. That is the defect this subsystem was rebuilt to close.

Both tokens and both context formats are wire values. They are mixed into key derivation and into every ciphertext at rest. Renaming a C# member is safe; changing a token or a context format is a breaking change that makes existing ciphertext unreadable.

Keys are derived, not used directly ​

The configured key is a master key. The key that actually encrypts is:

HKDF-SHA256(masterKey, salt = keyId, info = "gryd.secret.v2|" + purpose)

One key per keyId for the operator, one independent subkey per purpose in practice: leaking the MFA subkey exposes no client secret, and vice versa.

4. Errors an operator will see ​

CodeHTTPMeaningFix
SECRET_PROTECTION_UNAVAILABLE503No active key, or the envelope names a key this instance does not hold.Check the section is deployed; if a key was retired, put it back and finish the re-wrap first.
SECRET_PROTECTION_INVALID500The stored value is malformed, or failed its integrity check.Investigate the row — a tag failure means the ciphertext does not belong where it is stored.

Responses never carry the key id, the scope context, plaintext or ciphertext; the structured log carries the diagnosis.

A federated login whose client secret will not decrypt is refused as FederationException(ProviderUnavailable) rather than attempted with an empty secret — otherwise the operator would be debugging an authentication error at the customer's IdP instead of a key problem here.

5. Rotating a key ​

The rotation surface is admin:system only and deliberately not tenant-scoped: one key protects every tenant's secrets, so a rotation that could only see one tenant could never retire it.

GET  /api/v1/admin/security/secret-protection/status
POST /api/v1/admin/security/secret-protection/rewrap
  1. Generate a new key and add it to Keys, keeping the old one.
  2. Point ActiveKeyId at the new key and reload configuration — the key ring reads through IOptionsMonitor, so no restart is needed.
  3. POST .../rewrap with {"dryRun": true} and check the counts. (An empty body is a dry run.)
  4. POST .../rewrap with {"dryRun": false}. Repeat until totalRewrapped is 0 — the operation is idempotent.
  5. GET .../status until isFullyRotated is true and nothing is counted under the old key.
  6. Only then remove the old key from Keys.

Skipping step 5 is the mistake that produces the classic failure: a customer's first federated login of the day returns 503 because its secret still names a key that was deleted last night.

A re-wrap is also a tamper sweep ​

Each row is decrypted and re-encrypted under the scope rebuilt from the row itself. So the re-wrap cannot repair a mis-bound ciphertext — it can only refuse it. A non-zero failures count is a signal, not noise: that row's ciphertext does not belong where it is stored. Failures are reported by row id and never abort the run.

6. Envelope encryption with a KMS (optional) ​

With Scheme: "KmsEnvelope" each record gets its own AES-256 data key, wrapped by a KEK that never leaves the vault. The record itself is still sealed with the same associated data, so nothing about the domain separation changes — only the custody of the key does.

Reading always follows the scheme named in the ciphertext, so the migration is additive:

  1. Deploy with the KMS scheme available but Scheme: "KeyRing" — nothing changes.
  2. Provision the KEK and grant the deployment identity wrapKey/unwrapKey.
  3. Flip to Scheme: "KmsEnvelope". New secrets use the KMS; existing ones keep opening under kr.
  4. Run the re-wrap (§5) to convert the backlog.
  5. Remove the kr keys and their environment variables.

The data-key cache is in-process and must stay that way. Unwrapped data keys are cached in a private MemoryCache (bounded lifetime and entry count) so a federated login does not pay a network round trip. They must never reach the distributed cache: the federation descriptor deliberately carries the client secret to Redis still encrypted so a cache read is not a credential read — caching the key that opens it in the same Redis would undo that in one step.

Watch gryd_auth.secret_protection.kms.* and gryd_auth.secret_protection.data_key_cache.lookups (meter GrydAuth.Security.SecretProtection). An unwrap rate that tracks the login rate means the cache has regressed.

7. Migrating to this design (one-way) ​

The ResetSecretProtectionForV2 migration clears every secret protected under the old format. There is no reader for it by design, and the plaintext was never stored, so those values cannot be recovered.

Effect:

  • every MFA factor is set to Disabled and its owner must re-enrol;
  • every connection returns to Draft; an admin must re-enter the client secret and re-activate.

Deploy order

  1. Publish GrydAuth__Security__SecretProtection__* in every environment first.
  2. Deploy — the migration runs and clears the secrets.
  3. Admins re-enter each connection's client secret and re-activate it.
  4. Notify users that MFA re-enrolment is required.
  5. Verify: one federated login end to end, and one TOTP challenge after re-enrolment.

Before the deploy, list the tenants with enforceSso = true and check their break-glass accounts: while their connection sits in Draft, those users have no login path except break-glass.

Rollback restores the binary, not the secrets. Treat the deploy as one-way and rehearse it in staging with representative data.

Released under the MIT License.