- Jinja 28.2%
- Python 24.1%
- HTML 23.8%
- HCL 19.7%
- Shell 2.6%
- Other 1.6%
| ansible | ||
| compose | ||
| terraform | ||
| .gitignore | ||
| .sops.yaml | ||
| ansible.cfg | ||
| README.md | ||
Homelab — Infrastructure as Code de bout en bout
Infrastructure personnelle auto-hébergée, entièrement pilotée par code : provisionnement, configuration, sécurité et applications. Le projet documente une migration progressive depuis une installation OpenMediaVault/YunoHost bricolée à la main vers une infra reproductible, versionnée et automatisée.
Objectif du dépôt : servir à la fois d'infrastructure réelle en production et de démonstration technique — chaque brique reflète un vrai choix, avec ses compromis assumés plutôt que cachés.
Vue d'ensemble
Internet
│
▼
Traefik (reverse proxy, HTTPS via DNS-01/OVH)
│
├── Authelia (SSO, 2FA TOTP, OIDC pour certaines apps) ──> LLDAP (annuaire)
│
├── Forgejo (forge Git, public)
├── Grafana + Prometheus + Alertmanager (monitoring, alertes email)
├── Vaultwarden (gestionnaire de mots de passe)
├── Vikunja (gestion de tâches, SSO OIDC)
├── Portail applicatif maison (catalogue dynamique + admin LDAP)
├── Jellyfin (LXC, transcodage matériel iGPU)
└── Nextcloud (LXC, stockage utilisateurs sur le pool mergerfs, LDAP)
Semaphore (orchestrateur Ansible, secrets via SOPS, LAN uniquement)
Deux VM applicatives sur le même hyperviseur (dev / prod), même code, valeurs différentes — la VM dev sert de terrain de validation avant chaque bascule en production. Poste d'administration secondaire (ThinkPad X201) avec accès SSH par clé aux deux environnements.
Stack technique
| Couche | Outil | Pourquoi ce choix |
|---|---|---|
| Hyperviseur | Proxmox VE | Standard du homelab, gratuit, mature |
| IaC infra | OpenTofu (fork libre de Terraform) | Cohérence avec la philosophie open source du projet |
| Config management | Ansible | Rôle générique deploy_stack réutilisé pour tous les services |
| Secrets | SOPS + age | Secrets chiffrés versionnés dans Git, pas de service tiers à opérer |
| Stockage | SnapRAID + mergerfs | Protection contre la panne disque sans les risques de reconstruction du RAID 5 |
| Sauvegarde | vzdump (natif Proxmox) | Snapshots VM/LXC à chaud, planifiés |
| Reverse proxy | Traefik | Découverte automatique via labels Docker, HTTPS natif DNS-01 |
| Identité | LLDAP + Authelia | Annuaire léger + SSO (2FA, contrôle d'accès fin, OIDC natif pour les apps qui le supportent) |
| CI/CD | Semaphore UI | Alternative légère à AWX, jamais exposé publiquement |
| Forge Git | Forgejo | Fork communautaire de Gitea, gouvernance non lucrative |
| Monitoring | Prometheus + Grafana + Alertmanager | Observabilité VM/LXC/hôte, alertes email (SMART, disponibilité) |
| VPN torrent | PIA + thrnz/docker-wireguard-pia | Port forwarding natif, synchronisation automatique avec qBittorrent |
| Media | Jellyfin (LXC, iGPU) + Seerr (successeur de Jellyseerr) | Transcodage matériel Intel Quick Sync |
| Fichiers/cloud perso | Nextcloud (LXC) | Stockage utilisateurs adossé au pool mergerfs, LDAP natif |
| Tâches | Vikunja | SSO OIDC via Authelia |
Choix d'architecture notables
Docker en VM, pas en LXC, pour l'infra principale. Un LXC partage le kernel de l'hôte, ce qui introduit des restrictions de namespace incompatibles avec le fonctionnement normal de Docker (overlayfs, cgroups). La preuve en a été faite en pratique sur le LXC Nextcloud : nécessité d'activer nesting, conflit avec systemd-logind qui en résulte, décalage d'UID sur les bind mounts. La VM docker-host n'a jamais connu ces problèmes.
Authelia plutôt que Tinyauth. Le projet a démarré avec Tinyauth (minimaliste), puis migré vers Authelia pour le 2FA, le contrôle d'accès fin et l'OIDC natif — ce dernier permettant un vrai SSO sans double authentification pour les apps qui le supportent (Vikunja), sans la lourdeur d'Authentik.
Portail applicatif fait maison. Aucun outil existant ne propose un catalogue d'applications filtré par groupe LDAP avec la légèreté voulue. Le portail interroge dynamiquement l'API Traefik pour découvrir les services déployés et croise avec les groupes LDAP transmis par Authelia — aucune liste statique à maintenir. Une section d'administration permet la gestion des utilisateurs/groupes LLDAP via son API GraphQL.
Exposition minimale, même en interne. Chaque service est évalué individuellement — exposition publique uniquement si un vrai besoin existe, 2FA uniquement si l'enjeu de sécurité le justifie. Semaphore, qui a accès aux clés SSH et à l'exécution de playbooks sur toute l'infra, reste volontairement inaccessible depuis l'extérieur.
Services exclus du forward-auth. Vaultwarden et Jellyfin s'appuient sur leur propre authentification — leurs clients (extensions navigateur, apps mobiles) s'authentifient via API et ne peuvent pas suivre une redirection web classique.
Migration Jellyseerr → Seerr. Suite à la fusion du projet avec Overseerr début 2026, migration vers la nouvelle image (ghcr.io/seerr-team/seerr), avec ses changements de convention (plus de PUID/PGID, permissions gérées différemment).
Tests SMART robustes aux changements matériels. Le rôle Ansible lit dynamiquement /etc/snapraid.conf plutôt qu'une liste de disques codée en dur — il résout les vrais périphériques via findmnt, résistant à un renommage /dev/sdX (fréquent lors de l'ajout/retrait d'un disque) comme à l'ajout d'un nouveau disque au pool.
Sécurité
- HTTPS réel (Let's Encrypt via DNS-01/OVH) sur tous les services exposés, y compris en LAN
- Rate limiting Traefik (exclu des services à fort volume de requêtes : Semaphore, Vikunja, Authelia) + fail2ban
- Authelia : 2FA TOTP, contrôle d'accès par groupe, OIDC pour les apps compatibles
- Secrets chiffrés (SOPS/age) versionnés, jamais en clair dans Git
- Principe de moindre privilège systématique (tokens API scopés)
- LLDAP jamais exposé publiquement, même partiellement
- Sauvegardes VM/LXC régulières (vzdump), sur stockage dédié
- Alertes email automatiques sur dégradation SMART des disques du pool
Dette technique documentée
- Passthrough GPU et bind mount du LXC Jellyfin : Proxmox restreint ces opérations à
root@pam, incompatible avec le token API Terraform scopé. Gestes manuels documentés, à rejouer après toute recréation du conteneur. - LXC Nextcloud :
features.nestinget masquage desystemd-logindrequis pour que Docker fonctionne correctement ;user_account/clé SSH ne s'applique pas de façon fiable via Terraform sur les LXC, connexion enrootdocumentée comme contournement. - Export/montage NFS : configuré manuellement, pas encore en rôle Ansible.
- Création du réseau Docker
traefik_net: requis avant le tout premier déploiement Traefik sur une nouvelle VM, pas encore automatisé. - Statut par conteneur dans Grafana : aucune solution d'exporter mature identifiée à ce jour (écosystème fragmenté de petits projets peu maintenus) ; piste retenue pour plus tard : Uptime Kuma en parallèle, ou
docker-prometheus-exporter(calum4, encore jeune). pve_exporter: bug connu,pve_disk_usage_bytesreste à 0 pour les VM QEMU même avec guest agent actif ; contournement vianode_exporterinstallé directement dans les VM.
Structure du dépôt
homelab/
├── terraform/
│ ├── dev/ # VM de développement
│ └── prod/ # VM de production
├── ansible/
│ ├── inventories/{dev,prod}/
│ │ └── group_vars/all.sops.yml # secrets chiffrés
│ ├── roles/
│ │ ├── deploy_stack/ # rôle générique de déploiement Docker Compose
│ │ ├── monitoring_host/
│ │ ├── fail2ban/
│ │ ├── snapraid_maintenance/
│ │ ├── smart_test/ # lit snapraid.conf dynamiquement
│ │ └── vm_backup/
│ └── playbooks/
└── compose/
├── traefik/
├── lldap/
├── authelia/
├── forgejo/
├── semaphore/
├── monitoring/ # Prometheus, Grafana, Alertmanager
├── vaultwarden/
├── vikunja/
├── qbittorrent-vpn/ # PIA + thrnz/docker-wireguard-pia
└── portail/
Le rôle deploy_stack copie et templatise (SOPS/Jinja2) chaque stack de façon uniforme, avec rebuild/pull automatique à chaque déploiement.
Feuille de route
- Pipeline CI/CD : branches
dev/master+ webhooks Forgejo → Semaphore (piste Forgejo Actions identifiée, jamais implémentée) - Home Assistant (VM, HAOS) avec passthrough du dongle Zigbee/Z-Wave
- Déploiement de
node_exportersur la VM dev (actuellement seule la prod est couverte) - Automatisation des dernières étapes manuelles (réseau Docker initial, montage NFS)
- Solution fiable pour l'état par conteneur dans Grafana
Parcours du projet
Ce dépôt est le résultat d'une migration progressive, pas d'un design initial parfait : passage de YunoHost à une stack Docker maîtrisée, découverte et correction de nombreux bugs (indentation YAML, drift de version de provider Terraform, régression logicielle sur un client BitTorrent, incompatibilités LLDAP/OPAQUE, limitations Docker-en-LXC), arbitrages de sécurité documentés à chaque étape. L'historique Git reflète ce cheminement.
Homelab — End-to-end Infrastructure as Code
Jump back to the French version ↑
Self-hosted personal infrastructure, entirely driven by code: provisioning, configuration, security and applications. The project documents a progressive migration from a hand-tweaked OpenMediaVault/YunoHost setup to a reproducible, version-controlled, automated infrastructure.
Purpose of this repository: to serve both as real production infrastructure and as a technical showcase — every component reflects an actual decision, with its trade-offs stated rather than hidden.
Overview
Internet
│
▼
Traefik (reverse proxy, HTTPS via DNS-01/OVH)
│
├── Authelia (SSO, TOTP 2FA, OIDC for select apps) ──> LLDAP (directory)
│
├── Forgejo (Git forge, public)
├── Grafana + Prometheus + Alertmanager (monitoring, email alerts)
├── Vaultwarden (password manager)
├── Vikunja (task management, OIDC SSO)
├── Custom app portal (dynamic catalog + LDAP admin)
├── Jellyfin (LXC, hardware-accelerated transcoding)
└── Nextcloud (LXC, user storage on the mergerfs pool, LDAP)
Semaphore (Ansible orchestrator, SOPS-encrypted secrets, LAN only)
Two application VMs on the same hypervisor (dev / prod), same code, different values — the dev VM is the testing ground before any production rollout. A secondary admin laptop (ThinkPad X201) has key-based SSH access to both environments.
Tech stack
| Layer | Tool | Why this choice |
|---|---|---|
| Hypervisor | Proxmox VE | Homelab standard, free, mature |
| Infra IaC | OpenTofu (open-source Terraform fork) | Consistent with the project's open-source philosophy |
| Config management | Ansible | Generic deploy_stack role reused across every service |
| Secrets | SOPS + age | Encrypted secrets versioned in Git, no third-party service to run |
| Storage | SnapRAID + mergerfs | Disk failure protection without the risks of RAID 5 rebuilds |
| Backups | vzdump (native Proxmox) | Live VM/LXC snapshots, scheduled |
| Reverse proxy | Traefik | Automatic discovery via Docker labels, native DNS-01 HTTPS |
| Identity | LLDAP + Authelia | Lightweight directory + SSO (2FA, fine-grained access control, native OIDC for apps that support it) |
| CI/CD | Semaphore UI | Lightweight AWX alternative, never exposed publicly |
| Git forge | Forgejo | Community fork of Gitea, non-profit governance |
| Monitoring | Prometheus + Grafana + Alertmanager | VM/LXC/host observability, email alerts (SMART, availability) |
| Torrent VPN | PIA + thrnz/docker-wireguard-pia | Native port forwarding, automatic sync with qBittorrent |
| Media | Jellyfin (LXC, iGPU) + Seerr (Jellyseerr's successor) | Intel Quick Sync hardware transcoding |
| Personal cloud/files | Nextcloud (LXC) | User storage backed by the mergerfs pool, native LDAP |
| Tasks | Vikunja | OIDC SSO via Authelia |
Notable architecture decisions
Docker on a VM, not an LXC, for the core infrastructure. An LXC shares the host's kernel, which introduces namespace restrictions incompatible with Docker's normal operation (overlayfs, cgroups). This was proven in practice on the Nextcloud LXC: required enabling nesting, which in turn conflicted with systemd-logind, plus UID offset issues on bind mounts. The docker-host VM never ran into any of this.
Authelia over Tinyauth. The project started with Tinyauth (minimal), then moved to Authelia for 2FA, fine-grained access control, and native OIDC — the latter enabling true SSO without double authentication for apps that support it (Vikunja), without the overhead of Authentik.
A homegrown app portal. No existing tool offers an application catalog filtered by LDAP group with the lightness wanted here. The portal dynamically queries the Traefik API to discover deployed services and cross-references them with the LDAP groups forwarded by Authelia — no static list to maintain. An admin section handles LLDAP user/group management through its GraphQL API.
Minimal exposure, even internally. Every service is evaluated individually — public exposure only when there's a real need, 2FA only when the security stakes justify it. Semaphore, which holds SSH keys and can run playbooks across the whole infrastructure, stays deliberately unreachable from the outside.
Services excluded from forward-auth. Vaultwarden and Jellyfin rely on their own authentication — their clients (browser extensions, mobile apps) authenticate via API and can't follow a standard web redirect.
Jellyseerr → Seerr migration. Following the project's merger with Overseerr in early 2026, migrated to the new image (ghcr.io/seerr-team/seerr), along with its convention changes (no more PUID/PGID, permissions handled differently).
SMART tests resilient to hardware changes. The Ansible role reads /etc/snapraid.conf dynamically rather than a hardcoded disk list — it resolves real devices via findmnt, surviving /dev/sdX renaming (common when adding/removing a disk) as well as new disks being added to the pool.
Security
- Real HTTPS (Let's Encrypt via DNS-01/OVH) on every exposed service, LAN-only ones included
- Traefik rate limiting (excluded from high-request-volume services: Semaphore, Vikunja, Authelia) + fail2ban
- Authelia: TOTP 2FA, group-based access control, OIDC for compatible apps
- Encrypted secrets (SOPS/age) versioned, never in plaintext in Git
- Systematic least-privilege principle (scoped API tokens)
- LLDAP never publicly exposed, not even partially
- Regular VM/LXC backups (vzdump), on dedicated storage
- Automatic email alerts on SMART degradation of pool disks
Documented technical debt
- Jellyfin LXC GPU passthrough and bind mount: Proxmox restricts these operations to
root@pam, incompatible with the scoped Terraform API token. Manual steps documented, to be replayed after any container recreation. - Nextcloud LXC:
features.nestingand maskingsystemd-logindare required for Docker to work correctly;user_account/SSH key doesn't apply reliably via Terraform on LXCs, connecting asrootis documented as the workaround. - NFS export/mount: configured manually, not yet an Ansible role.
traefik_netDocker network creation: required before the very first Traefik deployment on a new VM, not yet automated.- Per-container status in Grafana: no mature exporter solution identified so far (fragmented ecosystem of small, poorly-maintained projects); candidates for later: Uptime Kuma alongside Grafana, or
docker-prometheus-exporter(calum4, still young). pve_exporter: known bug,pve_disk_usage_bytesstays at 0 for QEMU VMs even with the guest agent active; worked around vianode_exporterinstalled directly inside the VMs.
Repository structure
homelab/
├── terraform/
│ ├── dev/ # development VM
│ └── prod/ # production VM
├── ansible/
│ ├── inventories/{dev,prod}/
│ │ └── group_vars/all.sops.yml # encrypted secrets
│ ├── roles/
│ │ ├── deploy_stack/ # generic Docker Compose deployment role
│ │ ├── monitoring_host/
│ │ ├── fail2ban/
│ │ ├── snapraid_maintenance/
│ │ ├── smart_test/ # reads snapraid.conf dynamically
│ │ └── vm_backup/
│ └── playbooks/
└── compose/
├── traefik/
├── lldap/
├── authelia/
├── forgejo/
├── semaphore/
├── monitoring/ # Prometheus, Grafana, Alertmanager
├── vaultwarden/
├── vikunja/
├── qbittorrent-vpn/ # PIA + thrnz/docker-wireguard-pia
└── portail/
The deploy_stack role copies and templates (SOPS/Jinja2) every stack uniformly, with automatic rebuild/pull on each deployment.
Roadmap
- CI/CD pipeline:
dev/masterbranches + Forgejo → Semaphore webhooks (Forgejo Actions identified as the path forward, never implemented) - Home Assistant (VM, HAOS) with Zigbee/Z-Wave dongle passthrough
- Deploy
node_exporteron the dev VM (currently only prod is covered) - Automate the remaining manual steps (initial Docker network, NFS mount)
- A reliable solution for per-container status in Grafana
Project journey
This repository is the result of a progressive migration, not a perfect initial design: moving from YunoHost to a properly managed Docker stack, finding and fixing numerous bugs (YAML indentation, Terraform provider version drift, a software regression in a BitTorrent client, LLDAP/OPAQUE incompatibilities, Docker-on-LXC limitations), security trade-offs documented at every step. The Git history reflects that journey.