Skip to main content
Image coming soon

Internal PKI and Service Certificate Operations Evidence & Implementation Kit

$249.00
Adding to cart… The item has been added
Internal PKI and Service Certificate Operations · constrain the hierarchy, put identity in the certificate, attest before you sign, renew and reload, distribute trust in the right order, treat expiry as an outage class · Evidence & Implementation Kit
Run an internal certificate authority you can prove is operated, without a root whose custody nobody can reconstruct, an intermediate that can sign any name in the world, an identity that changes every time the scheduler moves the process, a renewal that succeeds while the process keeps serving the old certificate from memory, or an inventory that lists everything except the certificate that is about to take a service down.
Every control handed to you adopt-ready, from a root certificate authority defined as a written procedure with a stated key lifetime, a tamper-resistant device that never touches a routable network, a scripted ceremony with two authorised operators and an independent witness, and a succession point placed a full intermediate lifetime before expiry so the replacement is planned rather than discovered, through online issuing intermediates scoped one per environment and trust domain with lifetimes materially shorter than the root because replacing an intermediate is the only cheap recovery a hierarchy has, path length zero and name constraints on every intermediate proven by an out of scope signing request refused by each validating library, proxy and language runtime version actually deployed rather than by a command line verification tool that handles constraints differently, a single workload identity format carried in a URI subject alternative name with a trust domain and a hierarchical workload path that describes the service and its owning boundary rather than the host, node or address, so a redeploy or a region move leaves the identity unchanged and the common name carries no authorisation meaning at all, an identity registry fed by issuance events rather than by hand so a certificate cannot exist without a recorded owning team, owning system, business service and a contact route that reaches a person on call, weekly conformance measurement so authorisation policy matches a path segment instead of an enumerated allow list that grows with every deployment and that nobody prunes, attestation of every signing request against a property the platform itself asserts about the running workload so an automated issuer is not a signing service for anything that can reach it holding a widely readable bootstrap credential, the attestation source, the asserting component, its authenticating credential and the exact set of identities obtainable on its compromise written down per platform and signed off by the security function before production certificates flow, private keys generated inside the workload or its identity agent and delivered over a local socket or process-scoped channel rather than a shared volume that every other process, every backup and every crash dump can read, with a weekly scan of modes and owners rather than an assumption, renewal at a stated fraction of validity rather than a fixed number of days so shortening a lifetime does not break the schedule, randomised jitter so a fleet does not renew in one moment, an alert on the first failed attempt rather than at an expiry threshold, and a central distribution of remaining validity so the platform owner sees the shape of the estate rather than only its exceptions, verified reload proven by rotating a certificate and reading the new serial number off a live connection because the file on disk is exactly the thing that lies when a process read its certificate once at start-up, manual issuance refused in production with every unavoidable exception carrying a requester, an approver, compensating monitoring, a serial number and an expiry date, trust rotation run as three gated stages with new trust distributed and measured before anything presents it and old trust removed only after wire observation confirms nothing chains to it, mutual authentication enabled permissive first against an exit criterion written in measurements and a published date for strict mode so permissive does not become the destination, a four case negative test proving the server refuses an absent, an expired, an untrusted issuer and an out of policy client certificate rather than logging it and completing the handshake, an inventory built by active discovery of what every listening endpoint actually presents and reconciled against the authority so unmanaged certificates get owners, an expiry escalation ladder with a different named recipient at each rung and a monthly review that produces systemic fixes rather than chased certificates, and an honest revocation position that names what each library and proxy really checks, sets short lifetimes as the containment mechanism, and rehearses the break-glass reissue for a compromised intermediate before the day it is needed.
Ready in a weekend, not a quarter.

Here is the honest situation. Here is the honest situation. Internal PKI rarely fails because the cryptography is wrong. The libraries are mature, the defaults are mostly sensible, and standing up an issuing authority is a short piece of work that any competent platform team can complete in a sprint. It fails because the thing that was installed was never turned into an operation, and the gap between those two states only becomes visible on the day something expires or somebody asks who holds the root key. The first failure is the hierarchy. The root is described as offline, which turns out to mean it lives on a machine in a cupboard that is occasionally patched, or in a safe two people can open without a witness, and nobody can produce a record of what it has signed. Its lifetime was set by a default and never written into a plan, so the succession will be discovered rather than scheduled. Underneath it sits one issuing intermediate for everything, because that is the simplest thing to operate, which means production, test and any partner trust all rest on a single online key, and rotating it is an estate-wide event nobody will ever put in a change window. That intermediate is usually unconstrained, so it can sign a certificate for any name in the world including names the organization does not own, and the only boundary is the intention of whoever wrote the issuance profile. The second failure is the identity. Most internal certificates carry a name that describes where the workload is running rather than what it is, a host, a node, an address, sometimes a common name field that validators stopped treating as authoritative long ago. So the identity changes whenever the scheduler moves the process, no durable authorisation policy can be written against it, and the policy that exists becomes an enumerated allow list which grows with every deployment and which nobody has ever pruned. Nothing maps the identity to an owning team, so when a certificate approaches expiry the first activity is archaeology through repository history. The third failure is issuance. The moment issuance is automated the authority becomes an endpoint that mints identity on request, and in most estates the thing standing between an attacker with network reach and a valid production certificate is a shared bootstrap credential baked into an image and readable by anything on the node. That credential never appears on a risk register because it looks like plumbing. Where attestation does exist, nobody has written down what it actually trusts, so the question of what an attacker obtains by compromising a node agent gets researched during the incident. Private keys are generated centrally and written to a shared volume, which means every other process, every backup and every crash dump has a copy, and a certificate then proves only that something on that node had file access. The fourth failure is renewal, and specifically reload. Renewal gets automated first and reload gets forgotten, so a process that read its certificate once at start-up holds that material in memory for its whole life. The renewal succeeds, the file on disk is current, every dashboard is green, and the service keeps presenting an expired certificate until a client refuses the handshake. Nothing in the renewal system can see it, because the renewal system checks the file, and the file is the one thing that is telling the truth about a state the process is not in. Renewal is also scheduled too close to expiry, usually as a fixed number of days rather than a fraction of lifetime, so a single refused request during a maintenance window has no margin behind it. Alongside the automated estate sits a population of hand-installed certificates, each a small defensible exception on the day it was created, collectively the part of the estate with no owner and no renewal path. The fifth failure is trust distribution, which is where teams cause their own worst outage. If a service starts presenting a certificate from a new intermediate before every consumer holds the new anchor, every consumer refuses at the same instant, with no partial failure to warn anybody first. The safe order takes weeks of patience nobody budgeted, and the long tail is always the consumers nobody knew were consumers. Enforcement has the mirror problem. Permissive mode is the right first step, and then nothing forces the second step, because permissive looks perfectly healthy on every dashboard, so the estate carries the entire operating cost of certificates without the guarantee they were bought for. And where strict mode is claimed, it is frequently half true: the server requests a client certificate, logs the subject and completes the handshake regardless of whether the certificate is valid, absent, expired or signed by an unknown authority. Nobody runs the negative test, because the positive path works and no engineer volunteers to break a working service. The sixth failure is expiry itself, which is an outage class rather than an administrative task. Every organization that has had a certificate outage discovers afterwards that the certificate was not in the inventory, because the inventory was built from the issuing authority or from configuration and the certificate came from neither. The alerting is a single threshold into a shared channel that everybody has muted. And the revocation position is dishonest: a procedure exists, it has never been exercised, many client libraries do not check status by default, proxies soft-fail when the responder is unreachable, and a hard-fail configuration would make the responder a new single point of failure for every connection. So the organization believes a leaked key can be withdrawn on demand, when what actually happens is that lifetimes shrink, the issuing intermediate is rotated and the affected population is reissued, an operation nobody has rehearsed and everybody will be doing for the first time under pressure. Where teams fall short is predictable: a root nobody can account for, an intermediate that can sign anything, an identity that names a host, a registry that is a stale spreadsheet, issuance guarded by a credential in an image, keys on shared volumes, renewal without reload, trust rotated in the wrong order, mutual authentication that is half true, an inventory of what should exist, and a revocation plan that will not work on the day.

This Kit removes the guesswork. It is internal certificate authority and workload certificate operation written as adopt-ready controls you personalize in a weekend, with the evidence a head of platform engineering, a security function, an internal auditor or an incident review actually examines.

What you get, the moment you buy

18
Controls, adopt-ready. Every control, written so you personalize and apply it.
18
Evidence-they-examine checklists. For each control, exactly what a reviewer examines, plus where teams fall short, so you close the gap first.
1
Control Matrix, pre-built. Every control in a working spreadsheet, ready to record status, owner and evidence location.
1
Gap & Readiness Assessment. Score each control and the workbook returns your readiness as a single percentage, and exactly what to fix next.

Grounded in platform and infrastructure engineering practice as it is actually run inside estates where thousands of workloads authenticate to each other continuously and an expiry is an outage. Editable Word and Excel files. This is a practitioner method and it is honest about what certificates can and cannot give you, including where revocation does not work.

A certificate estate somebody operates, or an authority somebody installed and everyone now depends on
Internal PKI programmes are rarely stopped because the design was wrong. They lose credibility the morning a service fails a handshake, nobody can name the owner, the inventory does not contain the certificate, and the revocation procedure turns out not to work. This Kit builds the hierarchy, identity, issuance, renewal, trust distribution and expiry controls that keep those answers available before anybody has to ask.

What one control looks like

This is the opening control, where the hierarchy either becomes something you can prove custody of or stays a private key nobody can account for. All 18 are built to this depth.

HIER-1 Operate the root certificate authority offline as a written procedure with a stated lifetime and a named succession point, not as an assumption. CERTIFICATE AUTHORITY HIERARCHY AND TRUST BOUNDARIES
Put this control in place

Require [your organization name] to maintain a written root certificate authority procedure stating the root key lifetime, the algorithm and key size, the storage medium, the physical location and the named succession point at which a replacement root is generated. Require the root private key to reside in a hardware security module or equivalent tamper-resistant device that is never attached to a routable network, with the device serial number and the vault location recorded in the procedure. Require every root operation, including generation, intermediate signing and destruction, to run as a scripted ceremony with at least two authorised operators and one independent witness present, each named in the ceremony record. Require the ceremony record to capture the date, the participants, the exact commands run, the serial numbers of anything signed and the sealed-container numbers used to return the media to storage. Require the succession point to sit at least one full issuing intermediate lifetime before root expiry, so a replacement can be introduced without an emergency. Require the platform owner to review the procedure annually and after any change of authorised operator, recording the review date and reviewer.

Control note.

Involve physical security and whoever controls the safe before you write the ceremony, because vault access is usually the part the platform team has no authority to change. Rehearse the whole ceremony once with a disposable key so the first real run is not the first reading of the script.

Evidence a reviewer examines
  • The signed root certificate authority procedure, naming the root key lifetime, the algorithm, the storage device and the succession point.
  • The key ceremony record for root generation, listing the named operators, the independent witness and the commands run.
  • The hardware security module attestation or device inventory entry for the root key, showing serial number and vault location.
  • The sealed-container log covering every removal and return of the root media.
  • The signing record for each issuing intermediate, tying its serial number to a dated ceremony.
  • The dated annual review record of the root procedure, naming the reviewer.
Common finding they raise: Nobody can say who has held the root private key or what it has signed, so every certificate in the estate rests on an authority whose custody cannot be reconstructed.

Why this is not another template pack

  • The evidence is the point. A running certificate authority and a renewal client are not evidence. This tells you what a head of platform engineering, a security function, an internal auditor or an incident review examines and where teams fall short, for every control.
  • The hard specifics built in. An offline root defined as a witnessed ceremony with a recorded lifetime and a named succession point, issuing intermediates scoped one per environment and trust domain with lifetimes far shorter than the root, path length zero and name constraints proven by a negative test against the deployed proxy and runtime versions, a workload identity in a URI subject alternative name that survives a reschedule with the common name carrying no authorisation meaning, a registry fed by issuance so no certificate exists without an owner, weekly naming conformance so policy matches a segment rather than an allow list, attestation binding each request to a platform-asserted property with the blast radius written down and signed off per platform, keys generated in the workload and delivered over a process-scoped channel with weekly mode and owner scans, renewal at a fraction of lifetime with jitter and a first-failure alert, verified reload read off a live connection rather than off the file, manual issuance refused with owner and expiry on every exception, three gated trust rotation stages measured from the wire, permissive to strict with an exit criterion in measurements and a published date, a four case negative test proving refusal rather than a logged warning, an inventory from active discovery reconciled against the authority, an expiry ladder with a different named recipient per rung, and an honest revocation position with a rehearsed break-glass reissue are written into the controls, not left generic.
  • Built on real practice, not one person's opinion, grounded in how internal certificate authorities and workload identity actually hold together at scale and where that discipline usually breaks down.
  • It compounds. This work shares its shape with service authorisation, change management, incident response and supplier assurance, so it feeds your wider platform operating model.

Who buys this

Platform engineering and infrastructure leads, site reliability engineers, security engineers and the engineering managers accountable for internal certificate authorities and service to service authentication, who have to say who holds the root key and what it has signed, what an intermediate is allowed to issue, what a workload identity actually means, what happens when an issuing key leaks, whether a renewed certificate is really being served, and which certificate in the estate is closest to taking a service down. Whether you are standing up an internal authority that has to be defensible from the first issuance or formalising one that has quietly worked for a long time and been evidenced by nobody, you save weeks and walk in with your hierarchy, identity, issuance, renewal, trust distribution and expiry controls structured.

By the end of the weekend you will have
✓  An adopt-ready control for all 18 areas
✓  A completed control matrix
✓  The evidence a reviewer examines
✓  A written hierarchy and workload identity format
✓  A readiness percentage and a fix list
✓  The highest-risk gaps closed

Common questions

Is it really editable? Yes. Word and Excel files you own and adapt. No portal, no subscription.

Does it cover the whole practice? Yes. Certificate authority hierarchy and trust boundaries, workload identity in the certificate, attested issuance and private key custody, automated renewal and the reload problem, trust distribution and mutual TLS enforcement, and expiry as an outage class with revocation and break-glass each have their own controls with their own evidence.

Is this tied to one cloud, one orchestrator or one certificate authority product? No. The controls are principle-level, the hierarchy and constraint discipline, the identity format rules, the attestation and key custody requirements, the renewal and reload method, the trust distribution ordering, the enforcement proof and the expiry regime, so they apply whatever you run on and whatever issues your certificates.

Does it tell me what lifetimes to use? No, and it should not. Every lifetime, renewal fraction, threshold, conformance target and cadence in the Kit is a number your organization sets and records. What the Kit gives you is the method, the evidence and the discipline that makes your own numbers defensible.

Does this cover secrets, vaults and workload tokens as well? No. This Kit is certificates for workloads. Secrets management, vaults, workload identity federation and short-lived tokens in delivery pipelines, and key management inside a cloud key service, are separate subjects with their own Kits.

What if it is not for me? A 30-day money-back guarantee.

Do not let your next incident review be a root key nobody can account for, an inventory that did not contain the certificate, or a revocation procedure that was never going to work.
Every control is fast to adopt with the Kit. It is instant, and it is guaranteed.
Add it to your cart and be ready this weekend.

Instant digital download · 30-day money-back guarantee · The Art of Service Pty Ltd, GPO Box 2673, Brisbane QLD 4001 · support@theartofservice.com