# Building a Buffer-Inspired BIBE Lab from Floci to Full-Stack PR Previews

> **Snapshot:** 2026-09-21. Counts below are a point-in-time observation of the local Floci lab, not real AWS.
> **Scope:** A personal learning lab on one Ubuntu server. It is inspired by Buffer's public writing about Buffer Isolated Build Environments (BIBEs), but it is not Buffer's private implementation and it is not presented as production infrastructure.

This guide documents the path from an empty local AWS emulator to isolated, public, full-stack pull-request environments and a separate local-development stack. It includes the Floci AWS boundary, a fail-closed safety plane, Floci EKS backed by K3s, Kubernetes resource and security controls, Helm packaging, K3s network recovery, GitHub automation, Argo CD, DNS, TLS, Basic Auth, public and private repositories, OpenReview Studio, opt-in Floci SQS/KEDA scaling, and Floci ECR-backed Docker Compose.

No passwords, tokens, API secrets, private keys, or real AWS account identifiers are included.

## Contents

1. [What the lab proves](#1-what-the-lab-proves)
2. [Architecture and trust boundaries](#2-architecture-and-trust-boundaries)
3. [Host, tools, and final inventory](#3-host-tools-and-final-inventory)
4. [Install and isolate Floci](#4-install-and-isolate-floci)
5. [Configure the local AWS CLI boundary](#5-configure-the-local-aws-cli-boundary)
6. [Build the local AWS safety plane](#6-build-the-local-aws-safety-plane)
7. [Create Floci EKS and its K3s runtime](#7-create-floci-eks-and-its-k3s-runtime)
8. [Make the cluster safe and recoverable](#8-make-the-cluster-safe-and-recoverable)
9. [Prove namespace isolation manually](#9-prove-namespace-isolation-manually)
10. [Turn the prototype into a Helm release](#10-turn-the-prototype-into-a-helm-release)
11. [Automate create, verify, delete, and reboot recovery](#11-automate-create-verify-delete-and-reboot-recovery)
12. [Add GitHub and Argo CD GitOps](#12-add-github-and-argo-cd-gitops)
13. [Add public DNS, TLS, ingress, and authentication](#13-add-public-dns-tls-ingress-and-authentication)
14. [Support public and private GitHub repositories](#14-support-public-and-private-github-repositories)
15. [Deploy the OpenReview Studio full-stack preview](#15-deploy-the-openreview-studio-full-stack-preview)
16. [Use and remove a PR preview](#16-use-and-remove-a-pr-preview)
17. [Verification checklist](#17-verification-checklist)
18. [Failures that improved the design](#18-failures-that-improved-the-design)
19. [Cost, limitations, and next steps](#19-cost-limitations-and-next-steps)

---

## 1. What the lab proves

The final workflow is:

1. A developer opens a pull request in an approved GitHub repository.
2. A maintainer adds the `bibe-preview` label.
3. GitHub Actions builds immutable images tagged with the pull-request head SHA and publishes them to GHCR.
4. Argo CD's pull-request generator discovers the eligible PR.
5. Argo CD renders a trusted Helm chart from the GitOps repository.
6. Kubernetes creates one restricted namespace for that PR.
7. ExternalDNS creates a predictable hostname, ingress-nginx routes it, a wildcard certificate provides TLS, and host Nginx applies Basic Auth to the browser-facing route.
8. The reviewer uses the environment without sharing the developer's local machine.
9. When the PR is closed, merged, or loses the label, Argo CD prunes the Application and namespace; ExternalDNS removes the corresponding DNS records.

For OpenReview Studio, the preview is not one container. It declares Next.js, Fastify, an FFmpeg worker, PostgreSQL, Redis, MinIO, migration and bucket-initialization Jobs, ingress, policies, quotas, secrets, and persistent claims. Redis/BullMQ is the default job path; PR #1 opts into a Floci SQS queue and KEDA scales its worker between zero and two replicas. A separate Ubuntu Compose stack pulls matching prebuilt images from Floci ECR rather than real AWS ECR.

The lab also proves a separate but related safety track: local AWS control-plane state can be registered, frozen, evidenced, recovered with one-use tickets, cleaned exactly, and retained across a controlled Floci restart without touching real AWS.

## 2. Architecture and trust boundaries

The companion interactive diagrams are:

- [End-to-end architecture](/diagrams/buffer-bibe-architecture.html) — the platform overview.
- [Floci AWS service inventory](/diagrams/floci-aws-services-detailed.html) — the safety-plane services and observed SQS/ECR inventory, with ELBv2/ALB still absent.
- [Per-PR Pods and data stores](/diagrams/per-pr-pods-detailed.html) — individual OpenReview workloads; its Redis/BullMQ arrows show the default queue path.
- [SQS and ECR two-lane view](/diagrams/floci-sqs-ecr-two-lanes.html) — the current PR #1 SQS/KEDA path beside the separate local Compose/Floci ECR image path.

The design has three important boundaries.

### Boundary A: local AWS emulation

- AWS profile: `floci-local`
- endpoint: `http://127.0.0.1:4566`
- account: `000000000000`
- region: `ap-south-1`
- Floci is bound to loopback, not the public network.
- All AWS CLI examples in this track name the local profile and region explicitly.

This makes a real-AWS command visually different from a lab command. An endpoint check is performed before mutation-bearing drills.

### Boundary B: the Ubuntu BIBE platform

The Ubuntu host runs Floci, the K3s container created by Floci's EKS emulation, kubectl and Helm, Argo CD and ingress-nginx inside K3s, plus host Nginx, wildcard TLS, and local proxy services.

### Boundary C: each PR namespace

Each preview is restricted by restricted Pod Security admission labels, ResourceQuota, LimitRange, NetworkPolicy, ClusterIP-only Services, exact-SHA images, and controlled secret reflection for approved private repositories.

The pull request supplies application code. It does **not** supply trusted deployment policy. The Helm chart and platform controls remain in the GitOps repository's trusted main branch.

### One PR really means one namespace and its own pods

The isolation unit is not one shared preview Deployment. It is the complete namespace created for a particular repository and pull-request number.

For example, PR 4 and PR 5 can exist at the same time:

| Pull request | Argo CD Application | Kubernetes namespace | Public preview |
|---|---|---|---|
| PR 4 | repository-specific PR 4 Application | repository-specific PR 4 namespace | `https://<repo>-pr-4.git.example.com` |
| PR 5 | repository-specific PR 5 Application | repository-specific PR 5 namespace | `https://<repo>-pr-5.git.example.com` |

Each Application renders the same trusted chart with different PR values. Kubernetes therefore creates a separate set of Deployments, StatefulSets, Pods, Services, Ingresses, policies, Secrets and PVCs for each eligible PR. A commit to PR 4 rolls only PR 4. Closing PR 4 prunes only PR 4. PR 5 continues to run.

The shared components are the platform itself: the single K3s node, Argo CD, ingress-nginx, ExternalDNS, reflector, host Nginx and the wildcard certificate. Application workloads and state are namespace-scoped and are not shared between PRs.

#### Simple application PR

A simple repository such as the initial web preview gets:

- one Namespace;
- one Deployment with one application Pod;
- one ClusterIP Service;
- one Ingress;
- one ResourceQuota and one LimitRange;
- restricted Pod Security labels; and
- default-deny networking with an explicit ingress-controller allowance.

#### OpenReview full-stack PR

The default OpenReview chart declares six long-running workloads. PR #1 currently has five baseline Pods plus an FFmpeg worker Deployment scaled between zero and two Pods by KEDA:

1. Next.js web Pod;
2. Fastify API Pod;
3. FFmpeg worker Pod(s), one in the default BullMQ mode or zero to two in PR #1's SQS/KEDA mode;
4. PostgreSQL StatefulSet Pod;
5. Redis StatefulSet Pod; and
6. MinIO StatefulSet Pod.

It also gets two one-shot Job Pods that remain visible as `Succeeded` evidence:

1. a bucket-initialization Job; and
2. a Prisma migration and seed Job.

Around those Pods, the chart creates five ClusterIP Services, four Ingress objects, three PVCs, eleven NetworkPolicies, one ResourceQuota and one LimitRange. The Namespace is labelled for restricted Pod Security. This is why the OpenReview preview is a real multi-service per-PR environment rather than one generic demo Pod.

## 3. Host, tools, and final inventory

### Host assumptions

- Ubuntu 24.04-class host
- Docker Engine with rootful Docker available
- Nginx already serving the domain
- outbound internet access for images, GitHub, DNS API, and ACME
- a domain under your control
- sufficient CPU, memory, and disk for a single-node cluster and preview stack

The reference host had 16 CPU threads and enough memory for a bounded single-node lab. Do not copy the quotas upward merely because the host is larger; the quotas are part of the safety model.

### Client tools

The lab used AWS CLI v2, kubectl aligned to the actual K3s minor version, Helm 4, Docker and Docker Compose, Git and GitHub CLI, curl, jq, OpenSSL, and standard Linux networking tools.

Use official installation instructions. In particular, kubectl should stay within one minor version of the cluster control plane. The final cluster reported K3s `v1.34.1+k3s1`, and the host client was aligned to the Kubernetes 1.34 line.

### Final platform inventory

| Layer | Implemented state |
|---|---|
| Local AWS endpoint | Floci 2.0.1, loopback port 4566, persistent storage |
| Safety plane | IAM, Lambda, Step Functions, SNS, Scheduler, DynamoDB, S3 evidence |
| Kubernetes | Floci EKS facade backed by a persistent single-node K3s container |
| Package manager | Helm chart releases with lint, server-side dry-run, upgrade, test, and rollback history |
| GitOps | Argo CD, ApplicationSets, automatic sync, self-heal, and prune |
| Registries | GHCR for PR previews; three Floci ECR repositories for the separate local Compose stack |
| Edge | ingress-nginx, ExternalDNS, GoDaddy DNS, Let's Encrypt wildcard TLS, host Nginx |
| Access | Basic Auth for browser preview routes; SSO deliberately deferred |
| Repo support | Public and private GitHub repositories through separate least-privilege paths |
| Queue scaling | Floci SQS transcode queue and DLQ; KEDA bounds PR #1 workers at zero to two |
| Full-stack canary | OpenReview Studio web, API, worker, PostgreSQL, Redis, MinIO, Jobs and PVCs |
| Local development | Ubuntu Docker Compose, Floci ECR image copies, independent PostgreSQL/Redis/MinIO volumes |

### Exact Floci AWS service inventory

The AWS-shaped safety plane, Kubernetes preview platform and local-development stack serve different purposes. The SQS/ECR rows were checked against the running Floci endpoint on 2026-09-21; earlier safety-plane counts come from the retained lab baseline.

| AWS service | Used? | Retained lab resources | Purpose in this lab |
|---|---:|---|---|
| EKS | Yes | 1 cluster: `buffer-bibe-lab` | AWS-compatible cluster API that creates the local K3s runtime. It is not managed AWS EKS. |
| IAM | Yes | 6 service roles | Separates scheduler, Step Functions, adapters, controllers and cleanup permissions. |
| STS | Yes | local caller identities and role assumptions | Exercises AWS-style identities and service trust at the emulated boundary. |
| Lambda | Yes | 6 functions | Budget and TTL adapters, prepare/success/failure controllers, and exact cleanup. |
| Step Functions | Yes | 1 state machine: `buffer-lab-local-cleanup` | Orchestrates prepare, cleanup and terminal state with fail-closed transitions. |
| SNS | Yes | 1 topic: `buffer-lab-budget-stop-local` | Simulates the budget-stop event entering the cleanup plane. |
| EventBridge Scheduler | Yes | `buffer-lab-local` and `default` groups; zero active retained schedules | Represents the TTL trigger. Final drills removed the temporary active schedule. |
| DynamoDB | Yes | 1 table: `buffer-lab-runs` | Stores run identity, freeze, lock, execution and recovery-ticket state. |
| S3 | Yes | `buffer-lab-artifacts` plus Floci's Lambda task bucket | Stores evidence and exact registered objects used by cleanup drills. |
| SQS | **Yes, opt-in** | 2 queues: PR #1 transcode and DLQ | KEDA reads queue depth and bounds the FFmpeg worker at zero to two replicas. The default non-SQS path uses Redis/BullMQ. |
| ECR | **Yes, local only** | 3 repositories: OpenReview web, API and worker | Floci ECR stores exact-SHA image copies for the separate Ubuntu Compose stack. PR previews still pull GHCR images. |
| ELBv2 / ALB | **No** | 0 load balancers | Host Nginx and ingress-nginx perform public and in-cluster routing. No AWS ALB is provisioned. |
| AWS Budgets | **No Floci budget resource** | none | The SNS topic carries a simulated budget-stop envelope. The earlier real-AWS budget is a separate account control, not part of this local stack. |

This distinction matters: SQS and ECR were added for specific, separate purposes; they are not blanket replacements for Redis/BullMQ or GHCR across every preview. ELBv2/ALB remains absent. All these AWS-shaped resources live in Floci account `000000000000`, not Arun's real AWS account.

## 4. Install and isolate Floci

The Floci service lives in `/home/sam/floci-aws-lab`.

### 4.1 Create the directory

**Mutation — creates local directories:**

```bash
mkdir -p "/home/sam/floci-aws-lab/data"
cd "/home/sam/floci-aws-lab"
```

### 4.2 Create `compose.yaml`

```yaml
services:
  floci:
    image: floci/floci:2.0.1@sha256:4e451c39c7bb88e3cd4f87e8fc0c25d5b47695a51185d521e2241fa00486e8eb
    container_name: floci-aws-lab
    restart: unless-stopped
    user: "0:0"
    ports:
      - "127.0.0.1:4566:4566"
    environment:
      FLOCI_DEFAULT_REGION: ap-south-1
      FLOCI_DEFAULT_ACCOUNT_ID: "000000000000"
      FLOCI_BASE_URL: http://localhost:4566
      FLOCI_STORAGE_MODE: persistent
      FLOCI_STORAGE_PERSISTENT_PATH: /app/data
      FLOCI_SERVICES_S3_ENFORCE_AUTH: "true"
    volumes:
      - ./data:/app/data
      - /var/run/docker.sock:/var/run/docker.sock
```

Why each important setting exists:

- The image is pinned by tag and digest so an upstream change cannot silently alter the lab.
- `127.0.0.1:4566:4566` prevents LAN or internet access to the emulator.
- `persistent` plus the bind mount keeps emulator state across container restarts.
- the Docker socket is required for Floci EKS to create its backing K3s container.
- rootful Docker is required for this EKS path.
- S3 authentication is enforced so unsigned requests do not accidentally pass.

### 4.3 Start and verify Floci

**Mutation — starts a local container:**

```bash
docker compose up -d
```

**Read-only checks:**

```bash
docker inspect "floci-aws-lab" \
  --format 'Status={{.State.Status}} Restart={{.HostConfig.RestartPolicy.Name}}'

docker inspect "floci-aws-lab" \
  --format '{{range .Mounts}}{{println .Source "->" .Destination}}{{end}}'

curl --fail --show-error --max-time 10 "http://127.0.0.1:4566/"
```

Do not use `docker compose down -v` during routine work. The `-v` flag removes volumes and can destroy retained state.

## 5. Configure the local AWS CLI boundary

The Ubuntu package repository did not provide the required AWS CLI package, so AWS CLI v2 was installed from the official installer. Verify it first:

```bash
aws --version
```

### 5.1 Create a dedicated profile

Use dummy local credentials; never reuse real AWS credentials for an emulator profile.

**Mutation — writes local CLI configuration:**

```bash
aws configure set aws_access_key_id test --profile "floci-local"
aws configure set aws_secret_access_key test --profile "floci-local"
aws configure set region ap-south-1 --profile "floci-local"
aws configure set output json --profile "floci-local"
aws configure set endpoint_url http://127.0.0.1:4566 --profile "floci-local"
```

### 5.2 Prove that the profile is local

**Read-only:**

```bash
aws configure get endpoint_url --profile "floci-local"

aws sts get-caller-identity \
  --profile "floci-local" \
  --region "ap-south-1" \
  --output json \
  --no-cli-pager
```

Expected identity:

```json
{
  "Account": "000000000000",
  "Arn": "arn:aws:iam::000000000000:root"
}
```

This check matters more than convenience aliases. Before any mutation, confirm the profile endpoint, account, and region.

### 5.3 First service probes

**Read-only:**

```bash
aws s3api list-buckets \
  --profile "floci-local" \
  --region "ap-south-1" \
  --output json \
  --no-cli-pager

aws dynamodb list-tables \
  --profile "floci-local" \
  --region "ap-south-1" \
  --output json \
  --no-cli-pager
```

The first lab objects were an S3 artifacts bucket and a DynamoDB runs table. They later became the evidence and state stores for the safety plane.

## 6. Build the local AWS safety plane

The safety plane is not the PR runtime. It is a separate learning track that asks: can an automated environment be created, observed, frozen, recovered, and cleaned without guessing?

### 6.1 Components

- `buffer-lab-runs` DynamoDB table: active-run state, lock, freeze, execution records, and recovery tickets
- `buffer-lab-artifacts` S3 bucket: immutable evidence and drill artifacts
- SNS: notification fan-out
- Lambda: observer, adapters, execution controller, and cleanup handlers
- Step Functions: ordered, fail-closed workflow
- EventBridge Scheduler: time-bound cleanup trigger
- IAM roles and policies: explicit local trust and action boundaries

### 6.2 Build the deployment artifact

Source directory:

```text
/Users/shanthi/Documents/arun/buffer/lab-safety
```

On the machine containing the source:

```bash
python3 local-plane/build_local_plane_zip.py /tmp/buffer-local-plane.zip
```

The observer has local unit tests and strict input bounds. Its purpose is to validate an SNS-shaped event and record what would happen. It does not perform cleanup.

### 6.3 Deploy to Floci

**Mutation — creates local emulated AWS resources only:**

```bash
local-plane/deploy_floci.sh /tmp/buffer-local-plane.zip
```

The deployment script is intentionally provider-aware. It refuses to proceed if the profile, endpoint, account, or region does not match the Floci boundary.

### 6.4 Verify the plane

```bash
local-plane/verify_floci.sh
```

The tests progressed through four gates:

| Gate | Purpose |
|---|---|
| A | Validate local identity, deployed inventory, and positive observer path |
| B | Prove malformed payloads, oversized input, bad checksums, and unsafe registration fail closed |
| C | Run mutation-bearing cleanup drills only after evidence review and explicit authorization |
| D | Restart **only** the Floci container, verify persistent state, recover retained controls, and run a final positive canary |

### 6.5 Recovery-ticket design

The recovery mechanism uses an evidence-bound, single-use ticket. It prevents a generic “retry” from becoming an authorization to mutate arbitrary or changed state.

The rules are:

1. create evidence describing the exact blocked execution;
2. bind the recovery ticket to the evidence checksum and expected state;
3. consume it once;
4. reject reuse, mismatch, expiry, or missing evidence;
5. preserve `FREEZE=true` on unexpected state.

### 6.6 Final safety baseline

After Gate D, the retained local plane had persistent DynamoDB and S3 evidence, `FREEZE=true`, no active run, lock, or recovery ticket, zero active schedules, terminal execution records for the completed drills, and 117 local tests passing.

The exact counts are evidence from this lab snapshot, not universal Floci guarantees. Real AWS was not used by these drills.

## 7. Create Floci EKS and its K3s runtime

### 7.1 Install compatible clients

Install kubectl using the official checksum-verified instructions and keep it within one minor version of the actual cluster. Install Helm from an official release and verify the binary checksum.

```bash
kubectl version --client --output=yaml
helm version --short
```

### 7.2 Create a local IAM identity and EKS role

The lab created a local IAM user named `buffer-lab-kubectl` and an EKS service role in the emulated account. A separate profile, `floci-eks-admin`, used the same loopback endpoint and local-only credentials.

Before cluster creation:

```bash
aws configure get endpoint_url --profile "floci-eks-admin"

aws sts get-caller-identity \
  --profile "floci-eks-admin" \
  --region "ap-south-1" \
  --output json \
  --no-cli-pager
```

Expected account remains `000000000000`.

### 7.3 Create the cluster through the EKS API

**Mutation — creates a Floci-managed local K3s container:**

```bash
aws eks create-cluster \
  --name "buffer-bibe-lab" \
  --role-arn "arn:aws:iam::000000000000:role/buffer-lab-eks-role" \
  --resources-vpc-config "subnetIds=[],securityGroupIds=[]" \
  --kubernetes-version "1.37" \
  --profile "floci-eks-admin" \
  --region "ap-south-1" \
  --output json \
  --no-cli-pager
```

In this emulator, the EKS API request created a Docker container named `floci-eks-buffer-bibe-lab`. The AWS-facing version was metadata; the actual backing runtime reported K3s `v1.34.1+k3s1`.

### 7.4 Inspect both views

**Read-only:**

```bash
aws eks describe-cluster \
  --name "buffer-bibe-lab" \
  --profile "floci-eks-admin" \
  --region "ap-south-1" \
  --query 'cluster.{Name:name,Status:status,Endpoint:endpoint,Version:version}' \
  --output json \
  --no-cli-pager

docker inspect "floci-eks-buffer-bibe-lab" \
  --format 'Image={{.Config.Image}} Restart={{.HostConfig.RestartPolicy.Name}} Ports={{json .HostConfig.PortBindings}}'

docker exec "floci-eks-buffer-bibe-lab" kubectl get nodes -o wide
docker exec "floci-eks-buffer-bibe-lab" kubectl get pods -A
docker exec "floci-eks-buffer-bibe-lab" kubectl get --raw='/readyz?verbose'
```

The container used `rancher/k3s:latest`, mapped host port `6500` to container port `6443`, and stored K3s state in the Docker volume `floci-eks-buffer-bibe-lab`.

## 8. Make the cluster safe and recoverable

### 8.1 Fix the inotify ceiling

The first K3s container exited because the host allowed only 128 inotify instances. The durable host setting became:

```text
# /etc/sysctl.d/90-buffer-bibe-inotify.conf
# Allow the local Kubernetes lab alongside existing Docker services
fs.inotify.max_user_instances = 512
```

**Mutation — applies a host kernel setting:**

```bash
sudo sysctl -p "/etc/sysctl.d/90-buffer-bibe-inotify.conf"
```

Verify:

```bash
sysctl fs.inotify.max_user_instances
```

### 8.2 Generate and protect kubeconfig

The initial config was generated with:

```bash
mkdir -p "/home/sam/.kube"

aws eks update-kubeconfig \
  --name "buffer-bibe-lab" \
  --alias "floci-buffer-bibe" \
  --kubeconfig "/home/sam/.kube/floci-buffer-bibe.yaml" \
  --profile "floci-eks-admin" \
  --region "ap-south-1" \
  --endpoint-url "http://127.0.0.1:4566" \
  --no-cli-pager

chmod 600 "/home/sam/.kube/floci-buffer-bibe.yaml"
```

After reboot testing, the emulator-generated token path was not sufficiently recoverable. The active kubeconfig was rebuilt from the K3s client certificate and pointed to `https://localhost:6500`. The original token-based config was retained separately for evidence.

### 8.3 Block LAN access to the Kubernetes API

Docker-published ports can bypass assumptions made from UFW rules alone. The lab inserts a rule in `DOCKER-USER` and persists it with systemd.

```ini
# /etc/systemd/system/buffer-bibe-api-firewall.service
[Unit]
Description=Restrict Buffer BIBE Kubernetes API access
Requires=docker.service
After=docker.service
PartOf=docker.service

[Service]
Type=oneshot
ExecStart=/bin/sh -c '/usr/sbin/iptables -C DOCKER-USER -i eno2 -p tcp --dport 6443 -m conntrack --ctorigdstport 6500 --ctdir ORIGINAL -m comment --comment buffer-bibe-api-local-only -j DROP || /usr/sbin/iptables -I DOCKER-USER 1 -i eno2 -p tcp --dport 6443 -m conntrack --ctorigdstport 6500 --ctdir ORIGINAL -m comment --comment buffer-bibe-api-local-only -j DROP'
RemainAfterExit=yes

[Install]
WantedBy=docker.service
```

Enable it:

```bash
sudo systemd-analyze verify "/etc/systemd/system/buffer-bibe-api-firewall.service"
sudo systemctl daemon-reload
sudo systemctl enable --now buffer-bibe-api-firewall.service
```

Verify locally and from another LAN machine:

```bash
sudo iptables -L DOCKER-USER -n -v --line-numbers
kubectl --kubeconfig "/home/sam/.kube/floci-buffer-bibe.yaml" --context "floci-buffer-bibe" get nodes
```

From the Mac, `nc` to the Ubuntu LAN address on port 6500 timed out while local kubectl continued to work. That was the intended result.

## 9. Prove namespace isolation manually

Before automating PRs, the lab created namespaces `bibe-pr-001` and `bibe-pr-002` by hand. This made each control observable.

### 9.1 Restricted Pod Security

```bash
kubectl \
  --kubeconfig "/home/sam/.kube/floci-buffer-bibe.yaml" \
  --context "floci-buffer-bibe" \
  label namespace bibe-pr-001 \
  pod-security.kubernetes.io/enforce=restricted \
  pod-security.kubernetes.io/enforce-version=v1.34 \
  pod-security.kubernetes.io/audit=restricted \
  pod-security.kubernetes.io/audit-version=v1.34 \
  pod-security.kubernetes.io/warn=restricted \
  pod-security.kubernetes.io/warn-version=v1.34
```

An ordinary BusyBox pod was rejected because it did not set `allowPrivilegeEscalation=false`, drop all capabilities, run as non-root, and use an allowed seccomp profile. A corrected pod succeeded.

### 9.2 ResourceQuota and LimitRange

The baseline per preview was:

| Resource | Limit |
|---|---:|
| Pods | 8 |
| CPU requests | 1 core |
| CPU limits | 2 cores |
| Memory requests | 2 GiB |
| Memory limits | 4 GiB |
| Services | 10 |
| NodePort Services | 0 |
| LoadBalancer Services | 0 |
| Default container request | 100m CPU / 128Mi memory |
| Default container limit | 250m CPU / 256Mi memory |
| Maximum one-container limit | 1 CPU / 2Gi memory |

The tests deliberately attempted a 2-CPU container, a namespace request over 1 CPU, and a NodePort Service. All were rejected by the API server.

### 9.3 NetworkPolicy progression

The network tests were incremental:

1. confirm a labelled client could reach the web Service before isolation;
2. apply default-deny ingress and confirm the connection failed;
3. allow only a client with the approved labels and confirm access returned;
4. change the client label and confirm access failed again;
5. restore the label and confirm access returned;
6. add egress policies allowing DNS and the intended web destination only.

This is stronger evidence than checking that a NetworkPolicy object merely exists.

## 10. Turn the prototype into a Helm release

The chart lives at `/home/sam/buffer-bibe-lab/charts/preview-web` and packages ConfigMap content, Deployment, ClusterIP Service, default-deny and allow-client NetworkPolicies, security contexts and bounded resources, readiness and liveness probes, and a Helm test.

### 10.1 Lint and render

```bash
helm lint "/home/sam/buffer-bibe-lab/charts/preview-web"

helm template bibe-pr-002 \
  "/home/sam/buffer-bibe-lab/charts/preview-web" \
  --namespace "bibe-pr-002" \
  --kube-version "1.34.1" \
  --set-string 'preview.title=BIBE preview: PR 002'
```

### 10.2 Ask the real API server without persisting

```bash
helm template bibe-pr-002 \
  "/home/sam/buffer-bibe-lab/charts/preview-web" \
  --namespace "bibe-pr-002" \
  --kube-version "1.34.1" |
kubectl \
  --kubeconfig "/home/sam/.kube/floci-buffer-bibe.yaml" \
  --context "floci-buffer-bibe" \
  apply --dry-run=server -f -
```

`helm template` proves rendering. The server-side dry-run also tests Kubernetes schema, admission, policy, and quota behavior.

### 10.3 Install or upgrade atomically

```bash
helm upgrade --install bibe-pr-002 \
  "/home/sam/buffer-bibe-lab/charts/preview-web" \
  --kubeconfig "/home/sam/.kube/floci-buffer-bibe.yaml" \
  --kube-context "floci-buffer-bibe" \
  --namespace "bibe-pr-002" \
  --atomic \
  --wait \
  --timeout 120s
```

The Deployment contains a checksum annotation derived from the content ConfigMap. When content changes, the pod template changes and Kubernetes performs a rollout instead of leaving a pod mounted to stale content.

Inspect the release:

```bash
helm status bibe-pr-002 \
  --kubeconfig "/home/sam/.kube/floci-buffer-bibe.yaml" \
  --kube-context "floci-buffer-bibe" \
  --namespace "bibe-pr-002"

helm history bibe-pr-002 \
  --kubeconfig "/home/sam/.kube/floci-buffer-bibe.yaml" \
  --kube-context "floci-buffer-bibe" \
  --namespace "bibe-pr-002"
```

## 11. Automate create, verify, delete, and reboot recovery

The host lab includes:

```text
/home/sam/buffer-bibe-lab/scripts/create-preview.sh
/home/sam/buffer-bibe-lab/scripts/delete-preview.sh
/home/sam/buffer-bibe-lab/scripts/verify-lab.sh
/home/sam/buffer-bibe-lab/scripts/recover-after-reboot.sh
```

### Create

The create script validates the PR number, creates the namespace with restricted Pod Security labels, applies LimitRange and ResourceQuota, runs `helm upgrade --install --atomic --wait`, waits for rollout, and runs the Helm test.

### Delete

The delete script uninstalls the Helm release and removes the namespace. Namespace deletion is the cleanup unit, so namespaced Pods, Services, Secrets, Jobs, policies, and PVCs cannot be forgotten individually.

### Verify

The verification script checks the node, system pods, expected namespaces, releases, policies, quota, workloads, and service behavior. It ends only after the lab satisfies the baseline.

### Reboot recovery

Reboot exposed a non-obvious dependency: Docker reassigned the K3s container's default-bridge address while retained cluster state still expected the old one. The temporary repair used a bridge-address reservation and, when explicitly authorized, briefly restarted CI to recover address ordering. That workaround has now been replaced: K3s is pinned to `172.31.0.3` on the private `floci-aws-lab_default` network, with `node-ip: 172.31.0.3` and `flannel-iface: eth1` in `/etc/rancher/k3s/config.yaml`.

The updated `/home/sam/buffer-bibe-lab/scripts/recover-after-reboot.sh` checks the pinned endpoint, starts only the K3s container if needed, validates the advertised node IP, refreshes the certificate-based kubeconfig, and checks CI HTTP health. It does **not** reorder or restart CI. In a controlled test, K3s's default-bridge IP changed from `172.17.0.4` to `.3`, while the node stayed Ready at `172.31.0.3` and CI stayed healthy. The old reservation container is stopped and retained, not required for normal recovery. A pre-migration SQLite datastore backup is retained; neither it nor the live K3s volume should be deleted as a routine fix.

**Verification boundary:** this proves container-restart recovery, not a full Ubuntu or Docker-daemon reboot after the new pin. After the next host reboot, run the recovery script, `verify-lab.sh`, and the PR preview checks before calling host-restart persistence verified.

## 12. Add GitHub and Argo CD GitOps

The GitOps repository is `https://github.com/samarun/buffer-bibe-gitops`.

### 12.1 Remove the earlier GitLab path

The lab originally explored a local GitLab and runner path. It was later removed in favor of GitHub PR automation, while backup copies were retained rather than destructively deleting evidence.

### 12.2 Install Argo CD and ingress-nginx

Argo CD runs inside K3s and is exposed through a host-local proxy. The verified snapshot used Argo CD chart 10.9.1 / application v3.5.3, ingress-nginx, and Helm preview chart 0.2.0.

The bootstrap Application owns platform Applications. ApplicationSets discover PRs. Automated sync, self-heal, and prune maintain the desired state.

### 12.3 Trusted-main model

The critical security rule is:

```text
PR source code -> application image
trusted GitOps main -> Helm chart, policies, quotas, ingress, and lifecycle
```

An unreviewed pull request must not be able to replace its own NetworkPolicy, mount arbitrary host paths, create a privileged pod, or raise its quota by editing deployment configuration inside that PR.

### 12.4 GitHub Actions build contract

For an eligible PR, Actions checks out the exact PR head SHA, builds the required images, tags them with that immutable SHA, pushes them to GHCR, and exposes the SHA to the Argo CD/Helm values path.

Do not deploy `latest` for previews. The image reference is evidence connecting the running workload to the reviewed commit.

### 12.5 ApplicationSet lifecycle

The ApplicationSet PR generator filters for the `bibe-preview` label and creates a per-PR Argo CD Application. The chart owns its Namespace at an earlier sync wave, which fixed an early teardown problem where Argo CD removed workloads but left the namespace behind.

```text
label added -> Application created -> namespace and release reconciled
new commit -> exact-SHA values change -> deployment rolls forward
label removed / PR closed / PR merged -> Application pruned -> namespace deleted
```

## 13. Add public DNS, TLS, ingress, and authentication

### 13.1 Hostnames

The public pattern is:

```text
https://<repository-slug>-pr-<number>.git.arunsamuel.com
```

### 13.2 ExternalDNS and GoDaddy

ExternalDNS watches eligible Ingress resources and creates a CNAME for the preview hostname, a TXT ownership record so it can distinguish records it manages, and deletion of both records when the Ingress disappears.

GoDaddy API credentials are stored outside Git and injected through a Kubernetes Secret. They are never written into this guide or the GitOps repository.

### 13.3 Wildcard TLS

ACME DNS validation issued a wildcard certificate for:

```text
git.arunsamuel.com
*.git.arunsamuel.com
```

The certificate and private key live under `/etc/ssl/buffer-bibe` with restricted ownership and permissions. Renewal reloads Nginx only after a successful certificate install.

Let's Encrypt provides the certificate. It does not create DNS records, route traffic, or authenticate users; those are separate responsibilities.

### 13.4 Nginx and Basic Auth

Host Nginx terminates public TLS and proxies to local ports. Argo CD is available at `https://git.arunsamuel.com/`. Preview routes use a separate Basic Auth password file.

SSO/Keycloak was evaluated and deliberately deferred. For a personal lab, Basic Auth plus restricted access was enough to learn the infrastructure without introducing another stateful identity platform.

## 14. Support public and private GitHub repositories

The platform uses two different onboarding patterns.

### 14.1 Public repositories

- ApplicationSet can query public PR metadata.
- GHCR images may be public.
- no persistent GitHub repository credential is required inside the cluster.

### 14.2 Private repositories

Private onboarding separates discovery from image pull:

1. a GitHub App installed only on approved repositories provides read-only PR metadata;
2. a separate read-only package credential provides GHCR `read:packages` access;
3. the pull secret is stored centrally;
4. Reflector copies it only into namespaces carrying the approved label;
5. the chart attaches the pull secret to workloads that need the private image.

Selecting “all repositories” for the GitHub App is technically possible, but “only select repositories” is the safer default. It keeps the blast radius visible and makes onboarding intentional.

### 14.3 Repository adapter generator

The GitOps repository includes a generator:

```bash
python3 scripts/generate-repository-adapter.py --help
```

An adapter defines repository owner and name, public or private mode, image names, build workflow expectations, ApplicationSet template values, preview hostname slug, chart selection, and the secret-reflection label where required.

The public CuePilot Live and private Campus Connect pilots proved both paths, including complete teardown after merge.

## 15. Deploy the OpenReview Studio full-stack preview

OpenReview Studio was the full-stack canary because it exercises more than static HTTP.

### 15.1 Workloads and dependencies

```text
browser -> DNS/TLS/Nginx -> ingress-nginx
                              |-> /                     -> web Service    -> web Pod
                              |-> /api and /media       -> API Service    -> API Pod
                              `-> /originals /proxies   -> MinIO Service  -> MinIO Pod

web Pod -> API Pod -> PostgreSQL Pod
                   -> MinIO Pod
                   -> default: Redis/BullMQ -> worker Pod -> FFmpeg -> MinIO Pod
                   `-> PR #1: Floci SQS -> KEDA-scaled worker -> FFmpeg -> MinIO Pod
```

The web Pod serves the browser application and uses same-origin `/api` calls. The API Pod authenticates requests, reads and writes PostgreSQL, and signs object access. The default chart path creates BullMQ jobs in Redis. PR #1 instead has the `bibe-sqs` opt-in label: it sends transcode jobs to Floci SQS, and KEDA scales the FFmpeg worker Deployment according to queue depth. The worker writes proxies and thumbnails to MinIO and records status in PostgreSQL. Redis remains part of the OpenReview stack; SQS is not a blanket replacement for every Redis use.

The data services are individual StatefulSet Pods, not libraries inside the API container:

| Pod | Controller | Port | Per-PR persistent claim |
|---|---|---:|---:|
| PostgreSQL | StatefulSet | 5432 | 2Gi |
| Redis | StatefulSet | 6379 | 1Gi |
| MinIO | StatefulSet | 9000 | 8Gi |

The dedicated chart also creates two Job Pods, five Services, four Ingresses, eleven NetworkPolicies, one ResourceQuota, one LimitRange and runtime Secret references. The Job Pods initialize MinIO buckets and migrate/seed PostgreSQL, then stop successfully. They are expected to show `Completed`, not `Running`.

The observed lab snapshot contained both `bibe-openreview-canary` and `bibe-openreview-pr-1`. The manually retained canary keeps its default worker. The automated PR #1 namespace had five baseline running Pods, two completed Job Pods, and zero idle worker replicas because KEDA had scaled the worker down. When jobs arrive, PR #1 can have one or two worker Pods. The two namespaces do not share their databases or queues.

### 15.2 Bounds

The OpenReview preview baseline allows one web and API replica, plus one worker in default mode or zero to two workers in the opt-in SQS/KEDA mode. Worker concurrency is one per replica. The namespace retains a maximum of 10 pods, 6 Services, 3 PVCs and 12Gi requested storage, with zero NodePort or LoadBalancer Services and tightly controlled heavy-preview concurrency on this host.

The actual Namespace quota is:

| Resource | Hard limit |
|---|---:|
| Pods | 10 |
| CPU requests / limits | 2 / 6 |
| Memory requests / limits | 4Gi / 8Gi |
| PVC count | 3 |
| Requested storage | 12Gi |
| Services | 6 |
| NodePort Services | 0 |
| LoadBalancer Services | 0 |
| Jobs | 3 |

The LimitRange applies defaults of 50m CPU and 128Mi memory requests, 250m CPU and 256Mi memory limits, with per-container maxima of 2 CPU and 3Gi memory.

These constraints prevent a learning preview from quietly becoming an unbounded shared service.

### 15.3 Routing

| Path | Destination | Authentication behavior |
|---|---|---|
| `/` | Next.js web | Nginx Basic Auth |
| `/api` | Fastify API | application Bearer token |
| `/media` | API/media handler | application route rules |
| `/originals` | MinIO object path | presigned URL behavior |
| `/proxies` | MinIO object path | presigned URL behavior |

The frontend uses the relative URL `/api`. It must not embed a cluster-only service name or a temporary PR hostname during build.

### 15.4 Canary verification

The retained canary proved:

1. Argo CD Application `openreview-pr-1` became Synced and Healthy.
2. Web, API, and worker image references used the exact PR head SHA; an idle KEDA worker may have zero running Pods.
3. Browser login succeeded through the public TLS hostname.
4. The API reached PostgreSQL, Redis, and MinIO.
5. The default queue path created BullMQ work; the opt-in PR #1 path later used Floci SQS.
6. The worker ran FFmpeg with concurrency one per replica.
7. A playable proxy and a thumbnail were written to MinIO.
8. The browser could retrieve the media through the public route.
9. TLS and Basic Auth remained valid.

The canary PR stays open intentionally so the preview remains available for demonstration. That is a retention decision, not a cleanup failure.

### 15.5 Opt-in Floci SQS and KEDA

SQS and KEDA were added after the original BullMQ canary. The `bibe-sqs` label enables the Floci SQS path only for the eligible OpenReview PR. The observed Floci inventory has a transcode queue and a dead-letter queue for PR #1. KEDA reads queue depth through the private Floci endpoint and controls the worker Deployment with `minReplicaCount: 0` and `maxReplicaCount: 2`. These are emulator queues, not real AWS SQS queues.

**Read-only verification on Ubuntu:**

```bash
aws sts get-caller-identity --profile floci-local --region ap-south-1
aws sqs list-queues --profile floci-local --region ap-south-1 \
  --output json --no-cli-pager
kubectl --kubeconfig /home/sam/.kube/floci-buffer-bibe.yaml \
  --context floci-buffer-bibe -n bibe-openreview-pr-1 \
  get scaledobject,deployment
```

The deliberate test uploaded one video and observed worker scale-out from zero to one; two concurrent uploads reached the cap of two. All videos reached `READY`, queue and DLQ drained, and the worker returned to zero. That was a bounded local test, not a production throughput claim. Do not raise the two-worker cap without checking CPU, memory, storage, and competing previews on the shared Ubuntu host.

### 15.6 Hermes-style local stack with Floci ECR

The Ubuntu `openreview-bibe-local` Compose project is a **separate development environment**, not another PR namespace and not Buffer's private Hermes. It has a loopback web/API gateway on port 8093, and its own PostgreSQL, Redis and MinIO volumes. Three Floci-emulated ECR repositories hold copies of the exact-SHA OpenReview web, API and worker images. The PR preview continues to pull those images from GHCR; the local Compose stack pulls from the Floci ECR registry sidecar at `127.0.0.1:5100`.

**Read-only inventory and pull check on Ubuntu:**

```bash
aws ecr describe-repositories --profile floci-local --region ap-south-1 \
  --query 'repositories[].repositoryName' --output table --no-cli-pager
docker port floci-ecr-registry 5000/tcp
docker pull localhost:5100/000000000000/ap-south-1/bibe/openreview-web:sha-1d0d5a92abd4e6df333eec5758a2112286fa9b24
curl --fail --show-error http://127.0.0.1:8093/api/health/ready
```

The `localhost:5100/000000000000/ap-south-1/<repository>` form is Floci 2.0.1's local path-style registry transport. It is not a production ECR URI. The registry sidecar listens only on host loopback and persists images in a named Docker volume. The old optional local registry on port 5005 was stopped without deleting its volume. Application secrets and image references live in a mode-0600 untracked `.env.bibe-local`; never print or commit that file. A future image update should use a new `sha-<commit>` tag and verify its digest before changing Compose. Floci 2.0.1 records `IMMUTABLE` repository metadata but does **not** enforce tag immutability on push.

The local stack's web, API and worker were verified healthy after pulling from Floci ECR, and API readiness plus demo login passed. The registry also served an image after its own container restart. A full Ubuntu reboot of this new configuration has not yet been retested. OrbStack is optional on a Mac; the documented Ubuntu implementation uses Docker Engine.

## 16. Use and remove a PR preview

### Create or update

1. Open a pull request in an onboarded repository.
2. Wait for the required image-build workflow to succeed.
3. Add the `bibe-preview` label.
4. Watch Argo CD create the Application.
5. Open the predictable preview URL.
6. Push another commit to the PR and confirm the exact-SHA image changes and the deployment rolls forward.

### Merge

Use the repository's normal review policy. For a squash merge with GitHub CLI:

```bash
gh pr merge <PR_NUMBER> \
  --repo <OWNER>/<REPOSITORY> \
  --squash \
  --delete-branch
```

The repository name must be exact. A misspelling returns “Could not resolve to a Repository.”

### Teardown

Closing or merging the PR makes it ineligible. Argo CD should then remove the Application, namespace, Helm-managed workloads and PVCs, Ingress, and the ExternalDNS CNAME and TXT ownership record.

DNS and browser caches can make an old hostname appear alive briefly. Verify authoritative DNS, Argo CD Application absence, namespace absence, and a fresh uncached request before concluding teardown failed.

## 17. Verification checklist

### Floci boundary

- [ ] port 4566 listens on loopback only
- [ ] profile endpoint is `http://127.0.0.1:4566`
- [ ] account is `000000000000`
- [ ] persistent data mount exists
- [ ] container restart policy is `unless-stopped`

### Safety plane

- [ ] unit tests pass
- [ ] observer rejects malformed and oversized input without echoing payloads
- [ ] checksum and registration negatives fail closed
- [ ] recovery ticket is evidence-bound and single-use
- [ ] only the intended Floci container is restarted in Gate D
- [ ] retained state survives restart
- [ ] final canary performs exact cleanup

### Kubernetes platform

- [ ] node is Ready
- [ ] CoreDNS, local-path-provisioner, and metrics-server are Ready
- [ ] local API works on port 6500
- [ ] LAN access to port 6500 is blocked
- [ ] K3s advertises pinned `172.31.0.3` even when its default bridge IP changes
- [ ] after the next full host reboot, run the recovery and verification scripts; this new configuration has not yet passed that test
- [ ] firewall unit is enabled

### Preview controls

- [ ] restricted Pod Security enforced
- [ ] oversized and privileged pods rejected
- [ ] NodePort and LoadBalancer Services rejected
- [ ] unapproved client denied
- [ ] approved client allowed
- [ ] quota usage returns to zero after teardown

### GitOps and public edge

- [ ] Argo CD Application is Synced and Healthy
- [ ] running image tag equals PR head SHA
- [ ] wildcard TLS is valid
- [ ] Basic Auth protects browser routes
- [ ] API and presigned object routes keep their own auth semantics
- [ ] ExternalDNS creates and removes both CNAME and TXT records
- [ ] close or merge removes Application and namespace

### Opt-in queues and local images

- [ ] `floci-local` reports account `000000000000`, never the real AWS account
- [ ] PR #1's SQS transcode queue and DLQ exist only in Floci
- [ ] KEDA ScaledObject is Ready with worker bounds of zero to two
- [ ] the local registry binds only to `127.0.0.1:5100`
- [ ] all three Floci ECR repositories contain the expected exact-SHA image digests
- [ ] the separate Compose stack's web, API and worker are healthy, and API readiness passes

## 18. Failures that improved the design

| Failure | Root cause | Durable fix |
|---|---|---|
| AWS CLI `--name` missing in a multiline command | the backslash was not the final character or copied formatting changed it | put `\` as the final character with no trailing spaces; prefer one line while debugging |
| S3 PutObject returned `InvalidAccessKeyId` | request did not use the intended local profile/auth context | name `--profile floci-local` and region explicitly; verify endpoint first |
| K3s container exited | host inotify instance limit was 128 | persist `fs.inotify.max_user_instances=512` |
| `k3s kubectl` was an unknown nested command | the container entrypoint already exposed kubectl | use `docker exec <container> kubectl ...` |
| `localhost:6500` refused after reboot | K3s container was not running, or its retained state disagreed with a reassigned default-bridge IP | pin K3s to the Floci private-network address `172.31.0.3`; the recovery script verifies node IP without restarting CI. Full host reboot is still untested |
| UFW looked restrictive but Docker port was reachable | Docker forwarding path bypassed the expected UFW boundary | enforce the rule in `DOCKER-USER` and persist it with systemd |
| NetworkPolicy object existed but behavior was unclear | object inspection is not a connectivity test | run approved/unapproved client probes before and after labels change |
| Helm content updated without an automatic rollout | ConfigMap content was not part of pod-template identity | add a content checksum annotation to the Deployment |
| PR teardown left a namespace | namespace was created outside the Argo-owned lifecycle | make the chart own Namespace creation with an early sync wave |
| Nginx installer produced 400/502/401 checks | one test mixed proxy readiness and application authentication | split syntax, local-proxy, authenticated-app, and retrying route checks |
| Login showed “Failed to fetch” | frontend embedded a PR-specific API hostname that did not resolve | use same-origin relative `/api` and ingress rewrite |
| FFmpeg canary failed on a one-second video | thumbnail seek requested the frame at the clip boundary | seek inside the clip or select the first available frame |
| Merged PR URL still appeared | DNS/browser caching outlived the backend | verify Application, namespace, authoritative DNS, and uncached HTTP separately |

## 19. Cost, limitations, and next steps

### Cost

The local lab adds no EKS hourly fee because Floci and K3s run on an existing Ubuntu server. That does **not** mean the lab is free: the server, storage, electricity, backups, internet connection, domain, and operational time already have costs.

A future real-AWS validation must remain a separate project with a strict time-to-live, pre-approved resource types and regions, AWS Budgets notifications, automated cleanup based on registered exact state, residual inventory checks, and a fresh mutation authorization.

AWS Budgets is delayed and reactive; it is not a guaranteed hard spending cap or instant kill switch.

### Current limitations

- single Ubuntu host and single K3s node
- one pinned local K3s private-network address; full Ubuntu reboot of the new setup remains untested
- Floci ECR 2.0.1 does not enforce repository tag immutability on push
- Floci ECR is local emulation, not an authenticated production registry or real AWS ECR
- Basic Auth rather than SSO
- shared host failure domain
- no multi-zone availability
- no production backup/restore objective
- no formal SLO or on-call process
- one heavy OpenReview preview at a time
- wildcard certificate and DNS API operations are host-admin responsibilities

### Sensible next steps

1. add SBOM generation, image signing, and verification admission;
2. add policy-as-code such as Kyverno or Gatekeeper;
3. add a maximum preview age and abandoned-PR garbage collector;
4. add GitHub deployment/status comments with the preview URL and teardown result;
5. add observable backup and recovery drills for preview data where retention matters;
6. replace Basic Auth with SSO only when the identity lifecycle is worth operating;
7. run a short-lived real-cloud comparison under a separately approved budget and TTL;
8. compare local K3s behavior with managed EKS differences instead of assuming equivalence.

## References

- [Buffer's public BIBE article](https://peteremil.com/buffers-bibes/)
- [Floci project](https://github.com/floci-io/floci)
- [Floci ECR 2.0.1 behavior](https://github.com/floci-io/floci/blob/2.0.1/docs/services/ecr.md)
- [KEDA AWS SQS scaler](https://keda.sh/docs/2.20/scalers/aws-sqs/)
- [AWS CLI v2 installation](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html)
- [Install kubectl on Linux](https://kubernetes.io/docs/tasks/tools/install-kubectl-linux/)
- [Helm introduction](https://helm.sh/docs/intro/introduction/)
- [Argo CD ApplicationSet pull-request generator](https://argo-cd.readthedocs.io/en/stable/operator-manual/applicationset/Generators-Pull-Request/)
- [Kubernetes Pod Security Standards](https://kubernetes.io/docs/concepts/security/pod-security-standards/)
- [Kubernetes NetworkPolicy](https://kubernetes.io/docs/concepts/services-networking/network-policies/)
- [ExternalDNS](https://kubernetes-sigs.github.io/external-dns/)

---

## What this project demonstrates

This is evidence of hands-on learning across AWS-compatible APIs, failure containment, Docker, K3s/Kubernetes, Helm, GitHub Actions, GHCR, Argo CD, Floci SQS/KEDA, Floci ECR-backed local development, DNS automation, TLS, reverse proxies, registry isolation, stateful services, background media processing, teardown, and recovery. It should be described as a personal learning lab—not as ownership of Buffer's production infrastructure or production EKS operations.
