Floci → K3s → Helm → GitOps → PR preview

Building a Buffer-inspired BIBE lab from the ground up

This is the complete learning journey: a persistent Floci AWS emulator, a fail-closed safety plane, an EKS-created K3s cluster, Kubernetes and Helm controls, GitHub and Argo CD automation, public DNS and TLS, private repositories, a multi-container OpenReview preview, opt-in SQS/KEDA scaling, and Floci ECR-backed local development.

3 tracks

Safety + PR previews + local development

12 phases

Start-to-current build path

1 isolated namespace

And its own Pods per eligible PR

Exact SHA

Pinned image tags and verified digests

Auto-prune

Merge or close tears down

Track A

Local AWS safety plane

Floci, provider-bound CLI profiles, Lambda, Step Functions, SNS, Scheduler, DynamoDB, S3 evidence, freeze state, recovery tickets, exact cleanup, persistence, and Gate A–D drills.

Track B

Per-PR delivery platform

Floci EKS, K3s, kubectl, Helm, namespace controls, GitHub Actions, GHCR, Argo CD, ExternalDNS, wildcard TLS, Nginx auth, repository adapters, and the OpenReview stack with opt-in SQS/KEDA scaling.

Track C

Hermes-style local development

A separate Docker Compose stack on Ubuntu runs matching OpenReview web, API and FFmpeg images from Floci-emulated ECR, with its own PostgreSQL, Redis and MinIO volumes. This is a local learning loop, not Buffer’s Hermes or real AWS ECR.

Interactive complete system map

Everything from Floci to the browser

Switch themes, zoom, inspect components, follow the numbered flow, or export the diagram.

Floci service detail

AWS services actually used

The safety-plane map and inventory now include opt-in Floci SQS and local Floci ECR. ELBv2/ALB remains unused.

Per-PR detail

One PR expanded into its individual Pods

The OpenReview view separates web, API, worker, data-store and initialization Pods. Its Redis/BullMQ route is the default; PR #1 uses the SQS/KEDA overlay below.

Two distinct image and job paths

Floci SQS for the PR; Floci ECR for local Compose

The PR preview still pulls exact-SHA GHCR images. Its opt-in queue and worker scaling use Floci SQS and KEDA. Separately, the Ubuntu Compose stack pulls exact-SHA image copies from three Floci ECR repositories.

The isolation unit

Every eligible PR gets its own namespace and declared Pods

PR 4 and PR 5 can run at the same time because Argo CD creates two Applications, two namespaces, two Helm releases, two preview hostnames, and separate copies of every workload declared by the chart. A new commit rolls only that PR; closing it prunes only that PR.

1 running Pod

Simple application PR

  • ✓1 Namespace and 1 Deployment
  • ✓1 ClusterIP Service and 1 Ingress
  • ✓ResourceQuota, LimitRange, restricted Pod Security
  • ✓Default-deny networking with an explicit ingress allowance

5 baseline Pods + 0–2 workers + 2 completed Jobs

OpenReview full-stack PR

  • ✓Web, API, PostgreSQL, Redis and MinIO; PR #1 worker scales from 0 to 2
  • ✓Bucket-init and migration/seed Job Pods
  • ✓5 ClusterIP Services, 4 Ingresses, 3 PVCs and opt-in Floci SQS
  • ✓11 NetworkPolicies plus quota, limits, secrets, and Pod Security
PRNamespace
PR 4<repo>-pr-4
PR 5<repo>-pr-5

Evidence-backed AWS ledger

Used, emulated, and deliberately absent

This is the observed Floci inventory, not a generic AWS shopping list. SQS and ECR were added for distinct opt-in paths; ELBv2/ALB is still absent.

EKS

Used1 cluster

AWS-compatible API creates the local K3s runtime; it is not managed EKS.

IAM + STS

Used6 roles

Local identities, service trust, and least-privilege boundaries.

Lambda

Used6 functions

Budget and TTL adapters, three controllers, and exact cleanup.

Step Functions

Used1 state machine

Orchestrates fail-closed cleanup and terminal state.

SNS

Used1 topic

Carries the simulated budget-stop event.

EventBridge Scheduler

Used2 groups; 0 active schedules

Represents TTL expiry; temporary schedules are removed after drills.

DynamoDB

Used1 table

Stores runs, freeze, locks, execution, and recovery tickets.

S3

Used2 buckets

Stores evidence, exact registered objects, and Floci Lambda tasks.

SQS

Used2 queues

PR #1 transcode queue and dead-letter queue drive KEDA worker scaling; other previews may use Redis/BullMQ.

ECR

Used3 repositories

Floci ECR holds web, API and worker image copies for local Compose. PR previews still pull GHCR.

ELBv2 / ALB

Not used0 load balancers

Host Nginx and ingress-nginx route traffic.

Start-to-current build path

Twelve phases, each proven before the next layer

The visible preview was the last milestone. The platform began with provider isolation and failure containment, then moved upward through cluster recovery, namespace policy, packaging, and GitOps.

01

Local AWS boundary

Run Floci as an isolated, persistent AWS endpoint

I pinned Floci 2.0.1 by digest, bound port 4566 to loopback, mounted persistent storage, enabled S3 authentication, and created a dedicated floci-local AWS CLI profile with account 000000000000 in ap-south-1.

Evidence: The endpoint, account, region, image digest, mount, and restart policy are checked before mutation-bearing work.

02

Safety plane

Prove freeze, evidence, recovery, and exact cleanup

The emulated control plane uses IAM, Lambda, Step Functions, SNS, Scheduler, DynamoDB, and S3. Four gated drills tested malformed inputs, checksums, fail-closed registration, single-use recovery tickets, persistence, and exact cleanup.

Evidence: Gate D restarted only the Floci container, recovered retained controls, and finished with 117 local tests passing.

03

EKS to K3s

Create Kubernetes through the Floci EKS API

An emulated EKS cluster request created a persistent single-node K3s container. The API is local on port 6500, kubectl is aligned to the actual K3s minor version, and Helm manages application releases.

Evidence: Node readiness, system pods, and the API readyz endpoint were checked from both Docker and the host kubeconfig.

04

Host hardening

Fix runtime limits and make recovery repeatable

A low inotify ceiling caused the first cluster exit. A durable sysctl fix, a DOCKER-USER firewall rule, and certificate-based kubeconfig recovery hardened the host. Later, K3s moved from a dynamic Docker bridge address to a pinned private-network address.

Evidence: A controlled K3s restart changed its bridge IP while the node stayed Ready at 172.31.0.3; CI was not restarted. A full host reboot still needs retesting.

05

Kubernetes controls

Test one restricted namespace per preview

Manual PR namespaces established restricted Pod Security, ResourceQuota, LimitRange, ClusterIP-only Services, and default-deny NetworkPolicies before any GitHub automation was introduced.

Evidence: Oversized and insecure pods, NodePort Services, and unapproved clients were rejected; approved clients succeeded.

06

Helm lifecycle

Package, upgrade, test, and delete a preview as one release

The prototype became a Helm chart with health probes, security contexts, resource bounds, network policies, a test hook, and a ConfigMap checksum that triggers content rollouts.

Evidence: Lint, template, server-side dry-run, atomic install, upgrade history, rollback behavior, and namespace cleanup were exercised.

07

GitOps delivery

Connect GitHub Actions, GHCR, and Argo CD

GitHub Actions builds exact-SHA images. Argo CD ApplicationSets discover PRs carrying the bibe-preview label and reconcile a trusted chart from the GitOps main branch with automatic sync, self-heal, and prune.

Evidence: The chart owns the namespace, so closing or merging a PR removes both the Application and its complete environment.

08

Public edge

Add DNS, TLS, routing, and browser authentication

ExternalDNS manages GoDaddy CNAME and TXT ownership records, ACME DNS validation supplies wildcard TLS, ingress-nginx routes cluster traffic, and host Nginx applies Basic Auth to browser-facing preview paths.

Evidence: Preview DNS and TLS were tested through creation and teardown. SSO was considered and deliberately deferred.

09

Repository adapters

Onboard public and private repositories safely

Public repositories require no stored cluster GitHub credential. Private repositories use a read-only GitHub App for PR discovery and a separate read-only GHCR credential reflected only into approved namespaces.

Evidence: A public CuePilot pilot and private Campus Connect pilot both completed the create, update, and teardown lifecycle.

10

Full-stack canary

Run OpenReview Studio as a multi-container PR environment

Each eligible OpenReview PR gets its own Next.js, Fastify, FFmpeg worker, PostgreSQL, Redis and MinIO workloads, two Job Pods, three PVCs, Services, Ingresses, policies, quotas and runtime secrets.

Evidence: Login, upload, queue processing, FFmpeg proxy and thumbnail generation, public media access, exact-SHA images, TLS, and cleanup behavior passed.

11

Opt-in queue scaling

Use Floci SQS and KEDA for bounded FFmpeg workers

OpenReview PR #1 uses a Floci SQS transcode queue and dead-letter queue. KEDA reads queue depth and scales its FFmpeg worker Deployment between zero and two replicas; other previews can keep the Redis/BullMQ default.

Evidence: One upload scaled 0→1; two concurrent uploads reached the two-worker cap. All completed and the worker returned to zero.

12

Hermes-style local loop

Run matching images from Floci ECR in Docker Compose

A separate Ubuntu Compose stack now pulls prebuilt web, API and worker images from three Floci-emulated ECR repositories. Its PostgreSQL, Redis and MinIO volumes are independent of PR previews; the registry stays on loopback.

Evidence: Exact-SHA image pulls, container health, API readiness and login passed. No real AWS ECR was used; PR previews still pull GHCR images.

Technology inventory

What runs where

The lab is one system, but each layer has a distinct responsibility and trust boundary.

AWS-compatible control

Floci, IAM, Lambda, Step Functions, SNS, Scheduler, DynamoDB, S3, SQS and ECR

Kubernetes runtime

Floci EKS facade, K3s pinned to 172.31.0.3, kubectl, ingress-nginx and KEDA

Packaging and GitOps

Helm, GitHub Actions, GHCR for previews, Argo CD and ApplicationSets

Public edge

ExternalDNS, GoDaddy, ACME wildcard TLS, Nginx, Basic Auth

Preview guardrails

Pod Security, ResourceQuota, LimitRange, NetworkPolicy, exact-SHA images

OpenReview PR data plane

Next.js, Fastify, Floci SQS/KEDA for PR #1, FFmpeg, PostgreSQL, Redis and MinIO

Local development

Docker Compose pulls Floci ECR image copies; separate PostgreSQL, Redis and MinIO volumes

Guardrails

Safety is part of the architecture

Provider isolation

Loopback Floci endpoint, zero account ID, named local profiles, explicit region and preflight checks.

Mutation safety

Freeze state, evidence artifacts, checksums, one-use recovery tickets, exact-state cleanup, residual checks.

Namespace isolation

Restricted Pod Security, default-deny policy, bounded CPU/memory/storage, no NodePort or LoadBalancer.

Supply-chain identity

Preview workloads use PR-head-SHA image tags. Floci ECR stores SHA-named local copies, but its 2.0.1 registry does not enforce tag immutability.

Lifecycle ownership

Argo CD and the trusted chart own the namespace so prune removes the full environment.

Secret boundary

No secrets in Git; private repo discovery and private registry access use separate read-only credentials.

Failures turned into controls

The problems that changed the design

K3s exited during creation

Cause: The host allowed only 128 inotify instances.

Fix: Persist the host ceiling at 512 and verify it before recovery.

Kubernetes API disappeared after reboot

Cause: The retained K3s state and reassigned Docker bridge address diverged.

Fix: Pin K3s to 172.31.0.3 on the Floci private network; keep certificate kubeconfig recovery and fail-stop checks. Container restart passed; full reboot remains untested.

Firewall looked closed but Docker published the port

Cause: Docker forwarding did not match the expected UFW path.

Fix: Enforce the API restriction in DOCKER-USER and restore it through systemd.

Helm content changed without a rollout

Cause: A ConfigMap update did not change the pod template.

Fix: Hash the content into a Deployment annotation.

PR teardown left a namespace

Cause: The namespace lived outside Argo CD ownership.

Fix: Let the chart own Namespace creation at an early sync wave.

OpenReview login failed to fetch

Cause: The frontend embedded a PR-specific API hostname.

Fix: Use same-origin /api and ingress routing.

FFmpeg failed on the tiny canary

Cause: The thumbnail seek landed at the one-second clip boundary.

Fix: Seek inside the clip or take the first available frame.

Final verification

Evidence beyond “it deployed”

Positive paths and deliberate rejection tests cover the provider boundary, cluster, policy, release lifecycle, public edge, and application data path.

  • ✓Floci is loopback-only, persistent, and unmistakably separate from real AWS.
  • ✓Gate A–D safety drills fail closed and preserve evidence across a controlled restart.
  • ✓K3s stays Ready at pinned 172.31.0.3 when its default bridge address changes; full host reboot remains to be tested.
  • ✓Admission, quota, and network negative tests fail for the intended reason.
  • ✓Helm releases install atomically, roll on configuration change, pass tests, and delete cleanly.
  • ✓Argo CD reports eligible PR Applications Synced and Healthy with exact-SHA images.
  • ✓Public and private repository onboarding follow separate least-privilege paths.
  • ✓DNS, wildcard TLS, Basic Auth, API routing, object paths, and teardown were exercised.
  • ✓OpenReview login, upload, database, Floci SQS/KEDA scaling, FFmpeg, MinIO, and browser playback passed.
  • ✓The separate local Compose stack pulled exact-SHA images from Floci ECR and passed API readiness and login.

Cost and boundaries

A useful lab, not a claim of production equivalence

Floci and K3s avoid managed-EKS hourly charges by using existing hardware, but the server, power, storage, domain, internet, backups, and maintenance still cost money. The current design is single-node, uses Basic Auth, and has a shared host failure domain. KEDA caps the opt-in FFmpeg worker at two replicas. The pinned K3s address passed container-restart verification, but a full Ubuntu reboot is still untested. A real-AWS comparison remains a separate, short-lived exercise with TTL, budget alerts, exact cleanup, and residual checks.