Infrastructure & deployment

One Droplet to AWS

For two years, Garaad ran on a single DigitalOcean droplet. Here's what I moved to AWS, why, and what the decisions actually cost.

For two years, Garaad ran on a single DigitalOcean droplet. Django, Daphne, Celery, Redis, Nginx, and a Telegram media bridge — all on one box, with Postgres hosted separately on Supabase. It worked. Thousands of Somali developers learned on it. And every single deploy made me slightly nervous.

This is what I moved to AWS, why, and what the decisions actually cost.

Why leave

The droplet wasn't slow. It was undifferentiated.

One machine held the web tier, the worker tier, the cache, and the reverse proxy. If it died, everything died — and my recovery plan was "restore from a snapshot and hope." Deploys were me, SSH'd in, running git pull && docker compose up -d and watching logs scroll by. The database lived at a third provider, which meant two dashboards, two billing relationships, and a network hop between app and data that I couldn't see or measure.

The specific things that pushed me:

None of these were emergencies. That's exactly why it was the right time to move.

The shape of the new system

GitHub Actions ──OIDC──► IAM Role ──► ECR (build + push) │ └──SSM SendCommand──► EC2 (eu-north-1) │ ├─ backend (Django/gunicorn) ├─ daphne (ASGI/websockets) ├─ celery-worker ├─ celery-beat ├─ telegram-bridge (FastAPI) └─ redis │ ┌───────┴────────┐ ▼ ▼ RDS Postgres CloudWatch Logs

Deliberately not ECS or Kubernetes. More on that below.

Decision 1 — Managed Postgres instead of a third-party host

Supabase was fine as a database. The problem was that it was somewhere else. Connection strings pointed across the public internet, I had no visibility into slow queries from the same console as everything else, and my disaster-recovery story spanned two vendors.

Moving to RDS PostgreSQL in the same region as the app bought me automated backups with point-in-time recovery, encryption at rest, and — the part I underrated — the database being inside the same trust boundary as the application. The app talks to it over private networking with a security group, not over the open internet with a connection string.

It also surfaced a bug I'd been carrying for months. Deep in production_settings.py:

os.getenv('DATABASE_URL', 'postgresql://...@aws-0-us-east-1.pooler.supabase.com:5432/postgres')

A hardcoded fallback to the old database, sitting in version control. If the environment variable ever failed to load, production would silently connect to the previous provider and serve stale data — with a valid credential printed in the repo. I replaced the fallback with os.environ['DATABASE_URL'], so a missing variable crashes loudly instead of failing quietly into the wrong database.

Fail loudly beats fail silently. Migrations are excellent at exposing where you chose otherwise.

Decision 2 — No SSH. At all.

There is no SSH key for the production instance. Port 22 isn't open.

Deploys run through AWS Systems Manager. The EC2 instance carries an instance profile with AmazonSSMManagedInstanceCore, which lets the SSM Agent establish an outbound connection to AWS. Commands arrive over that channel. Nothing listens for inbound connections.

The deploy step is a single ssm send-command that pulls new images and recreates the containers:

- name: Deploy via SSM
  run: |
    aws ssm send-command \
      --instance-ids "${{ env.EC2_INSTANCE_ID }}" \
      --document-name "AWS-RunShellScript" \
      --parameters commands='[
        "cd /opt/garaad/garaad_backend",
        "git reset --hard origin/main",
        "docker compose -f docker-compose.prod.yml pull",
        "docker compose -f docker-compose.prod.yml up -d --no-deps redis",
        "docker compose -f docker-compose.prod.yml up -d --force-recreate --no-deps backend daphne celery-worker celery-beat telegram-bridge"
      ]'

The security win is real: a stolen laptop no longer grants shell access to production. The operational win is bigger than I expected — every command is logged in CloudTrail, attributed to an identity, with its output retrievable afterward. "Who ran what on the box, and when" stopped being a question I had to guess at.

Decision 3 — No long-lived AWS credentials anywhere

The repository contains zero AWS access keys. My own IAM user has none either.

GitHub Actions authenticates using OIDC federation. GitHub mints a short-lived signed token, AWS validates it against a registered identity provider, and returns a one-hour session. The trust policy is where the security actually lives:

{
  "Effect": "Allow",
  "Principal": { "Federated": "arn:aws:iam::<ACCOUNT_ID>:oidc-provider/token.actions.githubusercontent.com" },
  "Action": "sts:AssumeRoleWithWebIdentity",
  "Condition": {
    "StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com" },
    "StringLike":   { "token.actions.githubusercontent.com:sub": "repo:myorg/garaad_backend:ref:refs/heads/main" }
  }
}

That sub condition is the whole thing. Without it, the policy reads "any GitHub Actions workflow on earth may assume this role." With it, exactly one repository on exactly one branch can.

The permissions attached to that role are scoped just as tightly — push access to two named ECR repositories by ARN, and ssm:SendCommand restricted to one instance ID and one document. Not *. Not "all EC2 instances."

I checked this later with IAM Access Advisor: the role is allowed to touch two services and has used both, today. 100% of its granted permissions are in active use. The EC2 instance role, by contrast, showed 9 allowed services with only 5 ever accessed — because AWS-managed policies bundle in things you didn't ask for. CloudWatchAgentServerPolicy quietly grants X-Ray access I've never used.

Convenience-shaped policies are always broader than your workload. You can measure the gap instead of guessing at it.

Decision 4 — One EC2 box, not ECS or Kubernetes

This is the decision people push back on, so let me defend it.

Garaad's traffic fits comfortably on one instance. Moving to ECS or EKS would have added an orchestrator to learn, a networking model to debug, and a bill to justify — while solving a scaling problem I don't have yet. What I actually needed was reproducible builds, credential-free deploys, durable logs, and a database I don't operate myself. All four are solved above.

What I did take from the container world: images are built once in CI, pushed to ECR, and pulled by tag. The box never builds anything. A rollback is a tag change, not a rebuild.

Migrate the parts that hurt. Leave the parts that work.

What broke

Three things, all of them instructive:

  1. The Docker logging driver. I configured awslogs-stream-prefix in docker-compose.prod.yml. The installed Docker Engine rejected the option outright and refused to start the containers. Fixed by matching the config to the engine version actually on the box, not the one in the docs I was reading.
  2. Redis fell out of the deploy. The SSM command recreated backend, daphne, celery-worker, celery-beat, and telegram-bridge — but never redis. If Redis ever stopped, no automated deploy would bring it back, and Celery would silently lose its broker. Added a dedicated up -d --no-deps redis step, deliberately without --force-recreate so routine deploys don't drop broker connections for no reason.
  3. The hardcoded database fallback, above. The oldest bug, found by moving.

What isn't done

The S3 media bucket exists and the instance role is wired to it, but nothing writes there yet — user uploads still sit on the EC2 disk, which is the next real single point of failure. There's no load balancer, so there's no WAF and no second availability zone. And my own IAM user still carries AdministratorAccess directly, which is the pattern AWS spent a decade telling people to stop using.

Writing that list down is more useful than pretending the migration finished.

What I'd tell someone starting

Move the thing that is most expensive to operate, not the thing that is most fun to rebuild. For me that was the database and the deploy path — not the application itself, which barely changed.

And whatever you migrate, expect the move to find the bugs you've been carrying. Mine found three. That, on its own, was worth the weekend.

Building Garaad — a platform teaching Somali developers to build real software.

← More writing