Skip to content
version: 1.1.1

Platform Operations

What was delivered, and how it is run. This document covers the Azure DevOps control plane, the documentation release environments, and the runbook that takes a release candidate to production.

1. Azure DevOps control plane

All lifecycle phases, backlog items, and pipelines are consolidated in a single workspace.

  • Organisation: synkronyx
  • Project: EventRouteOptimiser
  • Repository: event-route-optimiser, a monorepo holding the edge API, the PWA front end, and the Bicep infrastructure

1.1 Service connections

Deployments run through isolated service connections, so role-based access control and network boundaries hold per environment.

Service connectionSubscriptionResource groupPurpose
sc-ero-devsub-sknx-ero-nonproductionrg-sknx-ero-developmentSandboxed development deployments
sc-ero-testsub-sknx-ero-nonproductionrg-sknx-ero-testAutomated integration testing
sc-ero-preprodsub-sknx-ero-nonproductionrg-sknx-ero-preproductionMirror staging environment
sc-ero-prodsub-sknx-ero-productionrg-sknx-ero-productionProduction systems

1.2 Deployment stages

Infrastructure deploys in ordered stages.

Stage 1, tenant scope. Provisions management groups and subscription aliases.

Terminal window
az login
az deployment tenant create --location uksouth --template-file platform/infrastructure/azure/main.bicep --parameters platform/infrastructure/azure/parameters/main.bicepparam

Stage 2, identity. Targets the newly created platform subscription and deploys the CIAM tenants.

Terminal window
$platformSubId = (az deployment tenant show --name main-bicep-deployment --query properties.outputs.platformSubscriptionId.value -o tsv)
az account set --subscription $platformSubId
az deployment sub create --location uksouth --template-file platform/infrastructure/azure/identity.bicep --parameters platform/infrastructure/azure/parameters/development.bicepparam

Stage 3, workloads. Deploys application workloads into the development, test, and production resource groups within their subscriptions.

1.3 Secret handling

  1. Environment credential files under settings/.env.*.local are excluded from version control.
  2. Pipelines retrieve secrets at run time from the Key Vault inside the relevant subscription boundary, or from mapped variable groups.
  3. Raw token values are never committed to any file under source control.
  4. Rotate CLOUDFLARE_API_TOKEN if it was ever exposed outside secret storage.
  5. Keep production, test, and development secret values separate when different Cloudflare accounts are used.

2. Documentation release environments

2.1 GitHub environments

Three environments carry the documentation release:

EnvironmentReviewer policy
docs-01-developmentNo reviewer requirement
docs-02-testOptional reviewer gate during release validation windows
docs-03-productionReviewer gate and wait timer for live release governance

2.2 Required secrets

Set these in all three environments.

SecretSourceNotes
CLOUDFLARE_API_TOKENCLOUDFLARE_API_TOKEN in settings/.env.localUsed by the Wrangler deploy action
CLOUDFLARE_ACCOUNT_IDMirrored from SKNX_CLOUDFLARE_ACCOUNT_ID, falling back to CLOUDFLARE_ACCOUNT_ID or TF_VAR_CLOUDFLARE_ACCOUNT_IDThe account id for the target Cloudflare account

2.3 Optional variables

Set these as environment variables when release branding and project names must match local settings. Each has a per-environment default, so omitting them is safe.

VariableDefault
SKNX_ERO_AZURE_COMPANY_NAMESynkronyx
SKNX_DOCS_PLATFORM_BRANDSynkronyx
SKNX_DOCS_GITHUB_REPO_URLThe repository URL
SKNX_DOCS_CORE_SITE_URLThe core site URL for the target lane
SKNX_DOCS_ERO_SITE_URLThe product site URL for the target lane
SKNX_DOCS_CORE_TITLESynkronyXr Core Docs
SKNX_DOCS_ERO_TITLEEvent Route Optimiser Docs
SKNX_DOCS_CORE_DESCRIPTIONSynkronyx shared platform standards and governance
SKNX_DOCS_ERO_DESCRIPTIONEvent Route Optimiser product documentation
CLOUDFLARE_CORE_PAGES_PROJECTThe core Pages project for the target lane
CLOUDFLARE_ERO_PAGES_PROJECTThe product Pages project for the target lane

Synchronise all of it from local settings with:

Terminal window
./platform/automation/powershell/sync-github-docs-environments.ps1 -Repository synkronyx/synkronyxr -Target all -WhatIf

Drop -WhatIf to apply, or pass -Target production, test, or development to scope the change. Add -SetRepositoryVariables to also write repository-level variables. This requires the GitHub CLI, valid authentication, and local settings files.

2.4 Published environments

This table is the single published index of documentation URLs.

LaneReached byEnvironmentCore siteProduct site
ProductionA docs-release-v* tagdocs-03-productiondocs.synkronyx.comero-docs.synkronyx.com
TestMerge to maindocs-02-testtest-docs.synkronyx.comtest-ero-docs.synkronyx.com
DevelopmentPer-branch and per-pull-request previewsdocs-01-developmentdev-docs.synkronyx.comdev-ero-docs.synkronyx.com

Cloudflare Pages projects, in the same order: sknx-docs and sknx-ero-docs, sknx-test-docs and sknx-test-ero-docs, sknx-dev-docs and sknx-dev-ero-docs. Each must return HTTP 200.

2.5 Routing

  1. Pushes to feature/** or user/** publish a per-branch preview URL. The shared development alias is never overwritten.
  2. Pull requests into main publish a per-pull-request preview URL, isolated from every other change in flight.
  3. Merging to main publishes to test.
  4. Tagging a commit on main with docs-release-v* publishes to production.

Merging and releasing are deliberately separate. A merge never reaches customers, so main can hold work that is not yet released, and releasing is an explicit decision about a specific commit rather than a side effect of merging a pull request.

Production is reached only by a tag. That means the commit validated in test is the exact commit published to production, with no intervening merge to introduce a difference.

2.6 Preview URLs

Branch previews follow this alias pattern:

https://<branch-alias>.sknx-dev-docs.pages.dev
https://<branch-alias>.sknx-dev-ero-docs.pages.dev

Cloudflare derives the alias from the branch name: lower-cased, non-alphanumeric characters replaced with hyphens, truncated to 28 characters. Pull request previews prefix it with pr-<number>-, so a preview for PR 59 on feature/enable-codespaces publishes as pr-59-feature-enable-codespa.

Always take the preview URL from the pull request comment. Do not reconstruct it, because truncation makes the alias ambiguous for long branch names. Previews are removed when their branch is deleted. Preview deployment is skipped for forked pull requests, because repository secrets are not exposed to forks.

2.7 Release tags

Versioning is GitVersion-driven. The workflow derives semVer from history and release tags, exports SKNX_DOCS_RELEASE_VERSION, and both sites display it in the sidebar label and the home page release callout. Commit messages may carry +semver: major, minor, patch, or none.

Tag patternEffect
docs-release-v*Both sites to production
docs-core-v*Core only to production
docs-ero-v*Product only to production
docs-test-release-v*, docs-test-core-v*, docs-test-ero-v*The same scopes to test
docs-dev-release-v*, docs-dev-core-v*, docs-dev-ero-v*The same scopes to development

3. Google identity licensing safe downgrade runbook

This runbook reduces Google Workspace licensing scope when Synkronyx only needs Cloud Identity controls for Google Cloud administration and federation prerequisites.

The goal is to lower recurring cost without breaking:

  1. Google Cloud organization administration.
  2. Group and administrator controls used by provisioning workflows.
  3. CIAM Google federation onboarding dependencies.
  4. Break-glass access for tenant and billing recovery.

3.1 Preconditions and hard safety rules

  1. At least two administrator accounts must be active before any license change.
  2. One break-glass administrator account must remain out of any federation dependency.
  3. The domain verification state must remain valid for every production domain.
  4. No downgrade step may start while incident response is active.
  5. Changes must run in a named maintenance window with rollback authority assigned.

3.2 Inventory and classification

Classify each current Google Workspace or Cloud Identity user into one category:

  1. Control plane administrator. Needs Google Cloud org, IAM, groups, and policy management.
  2. Service identity owner. Owns service accounts, workload identity federation, or billing automation.
  3. Collaboration user. Uses Gmail, Drive, Docs, or Meet.
  4. Dormant account. No required control plane or collaboration role.

Record each account in a change sheet with:

  • User identifier
  • Current license
  • Required capability
  • Target license
  • Migration owner
  • Rollback contact

3.3 Dependency checks

Before any license changes, confirm all dependencies:

  1. Google Cloud Organization access from at least two administrator accounts.
  2. Billing account administrator access from at least two administrator accounts.
  3. Group administration for all control-plane groups.
  4. Service account administration rights for all active automation projects.
  5. Existing CIAM onboarding fields remain collectible without Workspace-only features.

If any dependency fails, stop the downgrade and raise an open issue against the migration plan.

3.4 Pilot downgrade procedure

Run a pilot on one non-critical user first.

  1. Select one account from the collaboration or dormant category.
  2. Apply the target reduced license.
  3. Validate the account can still complete every required control-plane action in scope.
  4. Validate no required automation fails because of identity capability loss.
  5. Hold observation for one full business day before broader rollout.

If pilot validation fails, execute rollback immediately and document root cause in ERO Operations section 2 when the failure affects onboarding or secret workflows.

3.5 Wave rollout

After pilot success, run account changes in small waves.

  1. Roll out no more than 20 percent of in-scope accounts per wave.
  2. Validate control-plane, billing, and group administration after each wave.
  3. Record success and open issues before the next wave starts.
  4. Keep break-glass administrators unchanged until the final wave is complete.

3.6 Rollback runbook

Trigger rollback when any required control-plane function fails.

  1. Restore prior license assignments for the affected wave.
  2. Re-test organization, billing, IAM, and group administration paths.
  3. Confirm CIAM onboarding and Google federation inputs remain available.
  4. Freeze further downgrade waves.
  5. Open a post-incident task in Azure DevOps with a corrective action owner.

3.7 Completion criteria

The downgrade is complete only when:

  1. Required control-plane actions are validated for all in-scope administrators.
  2. Billing ownership and break-glass coverage are both confirmed.
  3. No automation failures are linked to license reduction.
  4. The final account and license map is stored in operations records.
  5. Cost deltas are recorded and approved by the change owner.

For the current Synkronyx operating model:

  1. Keep Microsoft 365 and Entra as the primary workforce identity substrate.
  2. Keep Google licensing focused on Cloud Identity and Google Cloud administration requirements.
  3. Use Google Workspace collaboration licenses only for users with active collaboration needs.
  4. Keep at least one dedicated break-glass administrator outside daily development workflows.

Release candidates do not create dedicated Pages projects. Candidate validation reuses the shared test environment, which is what a merge to main publishes to.

2.8 Edge cache

Every deploy purges the edge cache for the hosts it published. The purge runs unconditionally, because any deploy can change or remove a page.

The Cloudflare API token needs the Cache Purge permission on the documentation zone, granted through a zone-scoped policy. It does not appear under an account-scoped policy, which is the usual reason it seems to be missing. If the purge fails the workflow fails, because a release that leaves visitors on old content has not finished.

2.9 Stale published content, unresolved

Purging is a safety net, not a guarantee. There is an open defect where a published custom domain serves an old deployment indefinitely.

Observed on docs.synkronyx.com: a page deleted on 26 August was still served on 27 August, byte-identical to a deployment from 4 August, while the Pages project reported the current deployment as both latest and canonical.

Ruled out by measurement: DNS records, Pages custom domain binding, Worker routes, cache rules, _headers, Always Online, Internet Archive, and client-side caching. Attempted without effect: purge by hostname, purge_everything for the zone, Development Mode, and detaching and re-attaching the custom domain. Confirmed from an independent network, so it is not a local artefact.

Until this is resolved, treat a green release as evidence that the deployment succeeded, not that visitors see it. Verify explicitly:

Terminal window
Invoke-WebRequest https://docs.synkronyx.com/<removed-page>/ # expect 404

Validate published environments on their custom domains only. A pages.dev address is deployment infrastructure and can disagree with what visitors are served, which is exactly the condition this check exists to catch.

A 200 on the custom domain with a non-zero age header means the purge did not happen.

3. DNS is separate from publishing

Publishing content must never change infrastructure. The release workflow does not touch Terraform during a normal run.

dns_actionBehaviour
skipDefault. Publishes content only. Terraform never runs.
planReconciles state, prints the plan to the job summary, changes nothing.
applyRuns the same plan, then applies it if the safety checks pass.

Every push-driven and pull-request-driven run uses skip. TFSTATE_RESOURCE_GROUP, TFSTATE_STORAGE_ACCOUNT, and TFSTATE_CONTAINER are needed only for plan and apply.

This satisfies AG-IAC-002, which requires a reviewed plan before a production infrastructure apply. A content release carries no such plan, so it must not apply infrastructure.

3.1 Safety controls

The DNS path is guarded so that wrong or stale state cannot damage live records.

  1. Remote state is mandatory. There is no local or ephemeral fallback. Planning against empty state while live records exist is how records get silently overwritten, so the run fails instead.
  2. Existing records are imported before planning. A record present in Cloudflare but absent from state is imported first, so the plan reports an in-place update rather than a create over live configuration.
  3. Deletions and replacements abort the run. The plan is inspected as JSON and any delete or replace stops the workflow before apply. prevent_destroy on both records is the second line of defence.
  4. The plan is always published to the job summary, for plan and apply alike, so an apply is never unreviewable after the fact.
  5. The API token never reaches disk. Terraform reads it from the environment rather than a generated tfvars file.
  6. A split zone is rejected. Both hosts must sit in one Cloudflare zone, because the stack derives both record names from a single zone name.

Run plan first and read the summary. Only then re-run with apply.

4. Branch protection

Protect main, the only long-lived branch:

  1. Require a pull request before merge.
  2. Dismiss stale approvals when new commits are pushed.
  3. Require conversation resolution before merge.
  4. Block force-push and branch deletion.
  5. Permit merge commits only, so trunk history stays readable.

4.1 Required approvals

Required approvals are set to zero while Synkronyx has a single maintainer.

This is deliberate. GitHub does not permit self-approval, so a non-zero requirement on a single-maintainer repository cannot be satisfied and can only be met by bypassing the rule. Routine bypass is worse than no rule: AG-REL-001 classifies a manual bypass as a High finding, and a bypass that happens on every release makes a genuine bypass indistinguishable from noise in the audit trail.

The enforced gate is therefore the required status checks, which run on every pull request and do block. Raise the requirement to one the moment a second reviewer with write access exists.

4.2 Status checks

Because this repository mixes code and documentation, status checks need care:

  1. Do not make documentation-only workflows globally required checks. They do not run on every pull request, so a required check that never reports blocks every merge.
  2. Keep documentation workflows active and review their results when documentation changes.
  3. Use Universal PR Check and CAF and WAF Architecture Guard as the required checks on main.
  4. Add a further required check only when it runs on every pull request regardless of changed paths.

Because every change now reaches main through a single pull request, these checks are the only gate. They must stay trustworthy. A check that is allowed to fail routinely is worse than no check, because it trains everyone to merge past it.

Architecture review applies to any pull request touching architecture, infrastructure, identity, or governance. The review comment carries the <!-- architecture-guard-report --> marker, and backlog validation runs against its findings.

5. Release runbook

Releasing is a decision about a specific commit on main. It does not involve creating or merging a branch.

5.1 Land the work

  1. Branch from main, make the change, open one pull request into main.
  2. Review the per-pull-request preview URL published on the pull request.
  3. Merge once the required checks pass. Delete the branch.

The merge publishes to test automatically. Nothing has reached customers yet.

5.2 Validate on test

Validate https://test-docs.synkronyx.com and https://test-ero-docs.synkronyx.com. Soak for as long as the change warrants. Further merges to main republish test, so the lane always reflects the current trunk.

5.3 Release

  1. Tag the validated commit docs-release-vX.Y.Z, or use the scoped docs-core-vX.Y.Z or docs-ero-vX.Y.Z tags to publish one site.
  2. The docs-03-production environment gate applies, so the deployment waits for approval.
  3. Confirm https://docs.synkronyx.com and https://ero-docs.synkronyx.com.

The tagged commit is the exact commit that was validated on test. No merge happens between validation and release, so nothing can differ.

5.4 Roll back

Deploy the previous release tag. Production is whatever tag was last published, so rolling back is a deployment rather than a revert.

5.5 Patch a released version

Only needed when production must be fixed without shipping unreleased work that is already on main.

  1. Branch from the release tag, not from main.
  2. Apply the fix and tag it docs-release-vX.Y.Z+1.
  3. Merge the same fix into main so the trunk does not regress at the next release.
  4. Delete the branch.

This is the only circumstance in which a release branch is created, and it never outlives the patch.

5.6 Verification after configuration changes

When secrets or variables change, run the release workflow manually for core against development, then test, then run both against production, and confirm each deploys to the expected environment, that the production approval gate is enforced, and that both Pages projects receive updated artefacts.

6. Google Cloud organisation

The Google estate exists to hold OAuth clients for social federation and, later, workforce access to Google data. The identity plane model and project taxonomy are in Platform Architecture.

6.1 Licensing

Cloud Identity Free is the licence. It is sufficient to own a Google Cloud organisation, to create groups, and to federate sign-in with SAML, for up to 50 users.

Do not buy Google Workspace. Workspace exists to provide Gmail, Drive, and Meet. Synkronyx mail runs on Microsoft 365, and only one mail system can own the domain MX records. Cloud Identity Free is the edition without Gmail, so it does not compete for them.

Cloud Identity Premium adds context-aware access, advanced endpoint management, and automated provisioning to software as a service. None of that earns its cost at current headcount.

No artificial intelligence seat licences are required. Gemini Code Assist is a per-seat product and is not used, because GitHub Copilot is. Vertex AI and the Gemini API are billed per token against a project and carry no seat licence.

Automated user provisioning from Entra requires a Microsoft Entra ID P1 licence. At current headcount, creating Google accounts by hand is cheaper than the licence.

6.2 Federated administration

Staff sign in to Google with their Entra credential. The goal is one credential that the organisation manages, so no Synkronyx user holds a Google password for daily work.

Use Cloud Identity with SAML single sign-on, federated against the Synkronyx corporate Entra tenant. Do not use Workforce Identity Federation. Its users are not Google principals, hold no administration console identity, and require a credential file for gcloud. It solves a different problem at much higher complexity.

A Google identity object must exist for each person, because Google binds policy to a principal. A Google password does not. Those are separate concerns and only the first is unavoidable.

[email protected] holds Super Admin and is federated. A federated Super Admin can administer Google normally, so no separate account is needed for daily administration.

6.3 Break-glass account

[email protected] is the recovery account. It is Super Admin, is excluded from the single sign-on profile, holds a Google-managed password in the credential vault with a hardware security key, and is never used for daily work.

It exists for one failure mode. If the Entra signing certificate rotates or the SAML profile is misconfigured, every federated user is redirected to a broken identity provider, including every Super Admin. Repairing it requires signing in, and signing in is the broken step. The lockout is circular and Google will not readily resolve it.

Cloud Identity sign-up creates this account before single sign-on exists, so it is produced whether or not it is planned. The only decision is whether it is named and governed deliberately.

6.4 Bootstrap order

  1. Resolve any conflicting consumer Google account on a synkronyx.com address.
  2. Sign up for Cloud Identity Free and verify the domain by TXT record. The first account created becomes the break-glass administrator, so name it accordingly.
  3. Create the federated Super Admin account.
  4. Create the SAML single sign-on profile against the corporate Entra tenant, never against a CIAM tenant.
  5. Apply the profile to all users except the break-glass account.
  6. Prove federated sign-in in a private browser window before closing the working administrator session, so that a failed configuration can still be undone.
  7. Migrate any projects created outside the organisation into it. Projects do not join an organisation automatically when one is created later.

6.5 Project lifecycle

Project identifiers are immutable. Deletion is reversible for 30 days through gcloud projects undelete, but the identifier is destroyed permanently and can never be reused by anyone. Create the correctly named project first, then delete the incorrect one.

Project identifiers are 6 to 30 characters. Display names are 4 to 30 characters.

Projects are created by platform/automation/powershell/deploy/bootstrap-google-oauth.ps1, which parents them to google.organizationId and derives the identifier from the organisation prefix, the identity domain code, and the full tier word. Do not create projects by hand, because a manually created project is unparented by default and inherits no organisation policy.

6.6 Current estate

ProjectPurposeConsent screen
sknx-auth-nonproductionSynkronyx as relying party against Google APIsInternal
sknx-auth-productionSynkronyx as relying party against Google APIsInternal
sknx-ero-auth-nonproductionGoogle as identity provider into the ERO CIAM tenantExternal
sknx-ero-auth-productionGoogle as identity provider into the ERO CIAM tenantExternal

Synkronyx Life projects are created when that identity domain starts, not before. Creating projects ahead of the work that needs them is what produced the first round of misnamed projects.

7. Backlog

Backlog items are tracked as Epics, then Features, then User Stories, then Tasks, and are managed programmatically through the Azure DevOps MCP server. The structure required of them is in Platform Requirements.

The live board in Azure DevOps is the source of truth for work item state. This document deliberately holds no copy of it, because a duplicated status list goes stale the moment a work item moves.

7.1 Where working state belongs

The same rule governs every other place a list of open work could accumulate. Ranked from most durable to least:

  1. Azure DevOps. Open work, priority, assignment, and state. Anything a person could pick up and act on belongs here.
  2. The docs/ library. Decisions and the reasoning behind them. A decision does not go stale, which is what makes a document the right home for it.
  3. .github/agent-handover.md. Working state that has not yet reached the board, plus a pointer to what should. Nothing else.
  4. Assistant memory tooling. Nothing durable. It writes to editor storage on one machine, so it does not survive a machine change and cannot be reviewed, diffed, or corrected by anyone else.

The handover file is a bridge, not a destination. It exists because a session can end before its open issues are decomposed into work items, and losing them is worse than holding them briefly in a file. A list that lives there for long has the same defect as a status list in a document: it goes stale silently, because nothing forces it to be reconciled.

So a handover entry is a debt. Clear it by opening or closing the corresponding work item, then delete the entry. When the file holds nothing but the pointer to the board, it is doing its job.