<?xml version="1.0" encoding="utf-8"?>
<?xml-stylesheet href="/feeds.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:base="https://chameth.com/">
    <title>Chameth.com - posts like artisanal-docker-images, simple-backups-restic-hetzner, why-you-should-be-using-https but not migrating-from-github-to-forgejo</title>
    <subtitle>Personal homepage of Chris Smith</subtitle>
    <link href="https://chameth.com/feeds/posts/like/artisanal-docker-images,simple-backups-restic-hetzner,why-you-should-be-using-https/unlike/migrating-from-github-to-forgejo/" rel="self"/>
    <link href="https://chameth.com/"/>
    <icon>https://chameth.com/favicon.png</icon>
    <updated>2025-11-01T00:00:00Z</updated>
    <id>https://chameth.com/</id>
    <author>
        <name>Chris Smith</name>
    </author>
    <entry>
        <title>Thinking more about backups</title>
        <link href="https://chameth.com/thinking-more-about-backups/"/>
        <updated>2025-11-01T00:00:00Z</updated>
        <id>https://chameth.com/thinking-more-about-backups/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/thinking-more-about-backups/backblaze.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/thinking-more-about-backups/backblaze.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/thinking-more-about-backups/backblaze.png&#34; alt=&#34;The Backblaze logo: a stylised flame above the word Backblaze&#34; loading=&#34;lazy&#34; width=&#34;500&#34; height=&#34;320&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;The Backblaze logo&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Almost a year ago I wrote about &lt;a href=&#34;https://chameth.com/simple-backups-restic-hetzner/&#34;&gt;how I do backups with Restic and Hetzner&lt;/a&gt;.
That system has been ticking along well ever since, but recently I had some… thoughts. These backups are
all well and good if I accidentally delete a file, or a database gets corrupted, or something, but there are
two glaring issues:&lt;/p&gt;
&lt;p&gt;Firstly, I’m backing up my Hetzner server to Hetzner cloud storage. If something happens to Hetzner — or
my Hetzner account — then all my eggs go down with that basket. Obviously Hetzner are a big organisation
and aren’t likely to just vanish overnight, but I’m less confident about my account. Could a false abuse
report get it suspended? What if the UK passes
&lt;a href=&#34;https://www.legislation.gov.uk/ukpga/2023/50/contents&#34;&gt;even more dumb laws&lt;/a&gt; and Hetzner decide it’s easier
just to not do business with people here? This is the same sort of concern I have about Google accounts:
if you have half of your life in Google Drive and Google Mail, what happens if you comment on a YouTube
video, get flagged by an AI moderation process, and your account gets suspended? It’s probably not very
likely, but these are things my brain likes to dwell on.&lt;/p&gt;
&lt;p&gt;Secondly, the credentials to access the backups sit on each machine that is backed up. If someone malicious
gained access to the machine, they’d also have access to delete or tamper with all the backups. It feels
a little silly that the same attack could take down both the originals and the backups. There’s no way to
avoid that with Hetzner’s S3 implementation, as far as I can tell.&lt;/p&gt;
&lt;h3 id=&#34;exploring-options&#34;&gt;Exploring options&lt;/h3&gt;
&lt;p&gt;I toyed with the idea of making local copies of the backups, but the only way to avoid the same problems
would be to keep them offline and do a manual copy every now and then. I didn’t really want to do that,
and was concerned that if I did a monthly offline backup then I stood to lose up to a month of data in
the worst case.&lt;/p&gt;
&lt;p&gt;I then looked around at other S3 providers. &lt;a href=&#34;https://aws.amazon.com/s3/storage-classes/glacier/&#34;&gt;Amazon’s glacier offering&lt;/a&gt;
is tempting due to its very low storage costs, but you pay for that if you ever want to restore anything.
There are also lots of weird pricing edge cases around moving data between storage classes, minimum file
sizes, and so on. A much better option is &lt;a href=&#34;https://www.backblaze.com/cloud-storage&#34;&gt;Backblaze’s B2&lt;/a&gt; product.
Their pricing is much more straight-forward, and they have an interesting feature that’s particularly useful
in this case: &lt;a href=&#34;https://www.backblaze.com/blog/backblaze-b2-lifecycle-rules/&#34;&gt;lifecycle rules&lt;/a&gt;. Coupled with
the ability to create API keys that don’t have access to delete files (just “hide” them), this allows for
what’s effectively an append-only store.&lt;/p&gt;
&lt;p&gt;This works more-or-less out of the box with Restic. &lt;a href=&#34;https://pricey.uk/blog/restic-backups-without-delete/&#34;&gt;Joseph Price has a guide&lt;/a&gt;
that goes into the setup in a bit more depth. Basically, whenever Restic would delete a file (e.g. during
a “forget” or “prune” operation), it instead gets hidden and is only deleted when the B2 lifecycle rules
decide it should be. I’ve kept the existing Hetzner S3 backups for now, and just added an extra step to
the end of my script: a simple &lt;code&gt;restic copy&lt;/code&gt; and a &lt;code&gt;restic forget&lt;/code&gt;. B2 actually works out cheaper than the
Hetzner storage, as they don’t bill you for a minimum of 1TB storage; my current usage is around $3/month.
Not a bad price for some extra peace of mind!&lt;/p&gt;
</content>
    </entry>
    <entry>
        <title>Further Adventures in Music Organisation</title>
        <link href="https://chameth.com/further-adventures-in-music-organisation/"/>
        <updated>2025-09-21T00:00:00Z</updated>
        <id>https://chameth.com/further-adventures-in-music-organisation/</id>
        <content xml:lang="en" type="html">&lt;p&gt;I wrote before about how I’d &lt;a href=&#34;https://chameth.com/escaping-spotify-the-hard-way/&#34;&gt;dropped Spotify in favour of locally stored music&lt;/a&gt;,
but things have advanced a bit since. I had a few issues: Tauon would
occasionally manage to lose its database and along with it all my carefully
constructed playlists and song ratings&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:1&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;, and the experience on my phone was
not very fun.&lt;/p&gt;
&lt;p&gt;I had to manually sync the music by plugging my phone in to the
computer, and sometimes it just refused to mount the right partition. I don’t
think there’s really a good way to debug an Apple phone not behaving properly
when connected to a Linux desktop. Then I started wanting more than one playlist
synced&lt;sup id=&#34;fnref:2&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:2&#34; role=&#34;doc-noteref&#34;&gt;2&lt;/a&gt;&lt;/sup&gt;, and trying to find a way to make that work just broke me.&lt;/p&gt;
&lt;p&gt;I spent a while looking at different ways of hosting the music centrally.
&lt;a href=&#34;https://www.plex.tv/en-gb/plexamp/&#34;&gt;Plexamp&lt;/a&gt; gets lots of good reviews, and
I’ve used Plex a fair bit. I set about spinning up a Plex server, and just
could not get it working. The server ran fine, served the web interface, but
would neither associate with my account nor run standalone. The docs were
contradictory and there was very little useful logging. After a lot of
frustration, I stumbled across mentions that
&lt;a href=&#34;https://lowendbox.com/blog/plex-blocks-hetzner-in-move-against-piracy/&#34;&gt;Plex block running on Hetzner&lt;/a&gt;.
I assume that’s the cause of my issues, although I have no way to know for sure.
I use Hetzner for all my servers and other hosted services, but Plex have
decided I can’t run the self-hosted software that I have a lifetime subscription
for there. What the actual fuck?&lt;/p&gt;
&lt;p&gt;I could have probably worked around the arbitrary restriction, but I didn’t
want to throw more time down the drain. Instead, I set up
&lt;a href=&#34;https://www.navidrome.org/&#34;&gt;Navidrome&lt;/a&gt;, an open source music server. It
supports the Subsonic protocol, which means you can use a whole slew of
different clients with it (or even write your own). It also means there’s a
nice way to get data in and out of it programmatically, which I recall being
a bit of a fight with Plex.&lt;/p&gt;
&lt;!--more--&gt;
&lt;h3 id=&#34;syncing-and-organising&#34;&gt;Syncing and organising&lt;/h3&gt;
&lt;p&gt;All my music lived on my desktop, but now I wanted it on my server. At first
I just used &lt;code&gt;rsync&lt;/code&gt; to copy everything up. I’d maintain a local “master” copy,
and periodically shove the changes to the server&lt;sup id=&#34;fnref:3&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:3&#34; role=&#34;doc-noteref&#34;&gt;3&lt;/a&gt;&lt;/sup&gt;. That quickly got old,
so I set up &lt;a href=&#34;https://syncthing.net/&#34;&gt;Syncthing&lt;/a&gt; to keep the two folders in
sync. That worked… for a while.&lt;/p&gt;
&lt;p&gt;Concurrently, I was looking at improving the organisation of my music. It
was mostly organised by some Go programs I’d thrown together to automate the
importing, with no real validation of metadata or anything else. Some albums
got split up into multiple folders because the tracks had different artists
(and didn’t have an album artist set); some artists ended up with multiple
folders with slight spelling or case variations. It was upsetting.&lt;/p&gt;
&lt;p&gt;I’d used &lt;a href=&#34;https://picard.musicbrainz.org/&#34;&gt;MusicBrainz Picard&lt;/a&gt; before to fix
some issues, but I didn’t really like the UI, especially when trying to do
bulk actions. The main alternative is &lt;a href=&#34;https://beets.io/&#34;&gt;beets&lt;/a&gt;, which
describes itself as “the music geek’s media organiser”, and is entirely
command line. I was intrigued.&lt;/p&gt;
&lt;p&gt;I started playing around with beets, and quickly noticed a problem. Every time
I changed a metadata tag in a music file, Syncthing had to upload the entire
file. I was frequently changing tags across the whole library as I got beets
set up how I wanted, and Syncthing was handling it by uploading gigabytes of
files for every minor change. I was already a bit unhappy with having the
content duplicated in two places, so now I was using a command line organiser
I figured I could just get rid of my local copy and move everything to the
server.&lt;/p&gt;
&lt;p&gt;I set up a Docker container for beets, and then wrote some incredibly hacky
shell scripts so I could run &lt;code&gt;beet&lt;/code&gt; locally on my desktop and it would SSH
to the server over Tailscale and exec into the container, passing the arguments
along. Then I did one final sync of the music library, got rid of Syncthing,
double checked I had a backup and deleted all of my local music.&lt;/p&gt;
&lt;h3 id=&#34;bears-beets-battlestar-galactica&#34;&gt;Bears, Beets, Battlestar Galactica&lt;/h3&gt;
&lt;p&gt;Beets is amazing. It has a vast array of plugins that can do almost anything
you could want with a music library, and the command line workflow works really
well for me. It’s very well documented, and all the individual parts are
pleasingly simple and easy to understand.&lt;/p&gt;
&lt;p&gt;Here’s what it looks like when importing some new music:&lt;/p&gt;
&lt;figure class=&#34;image full&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/further-adventures-in-music-organisation/beets-importing.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/further-adventures-in-music-organisation/beets-importing.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/further-adventures-in-music-organisation/beets-importing.png&#34; alt=&#34;Screenshot of beets output when importing an album. It shows the source folder in blue, the matched metadata information in white with green highlights, the MusicBrainz URL, then a listing of all tracks&#34; loading=&#34;lazy&#34; width=&#34;782&#34; height=&#34;377&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;Beets importing an album&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;If it doesn’t get a perfect match then it shows the closest matches, and
summarises what’s different about them (missing tracks, different names,
etc), and lets you decide what to do. Aside from the metadata matching, it’s
doing a lot of things under the hood that aren’t necessarily apparent. It:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Looks up the genre from Last.fm&lt;/li&gt;
&lt;li&gt;Analyses all the tracks and writes ReplayGain metadata to them&lt;/li&gt;
&lt;li&gt;Fetches album art&lt;/li&gt;
&lt;li&gt;Checks all the files to make sure they’re actually playable&lt;/li&gt;
&lt;li&gt;Scrubs any existing metadata tags&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;It also maintains its own database, and you can query against it. The query
language is both simple and quite powerful, like the rest of beets. As an
example:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ beet ls artist:&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;Linkin Park&amp;#34;&lt;/span&gt; year:2025 length:2:00..2:30 
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Linkin Park - From Zero - Casualty
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Almost all the other commands also let you use the query syntax, which makes
the whole tool really powerful. Want to delete all country tracks that you
added last month? No problem. Redownload all the album art for a certain
band? Easy. You can even do smart playlists using the filters.&lt;/p&gt;
&lt;h3 id=&#34;actually-playing-music&#34;&gt;Actually playing music&lt;/h3&gt;
&lt;p&gt;With all this organisation, I’ve not actually mentioned one tiny detail: how
I actually play music now. Navidrome has a web UI, which is perfectly usable,
but I don’t really want my web browser involved as it makes balancing sound
levels tricky, and getting media keys working is a pain. Thankfully there
are loads of Subsonic clients. I’ve tried several, but eventually settled on
&lt;a href=&#34;https://github.com/jeffvli/feishin&#34;&gt;Feishin&lt;/a&gt; on the desktop. It’s very pretty:&lt;/p&gt;
&lt;figure class=&#34;image full&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/further-adventures-in-music-organisation/feishin.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/further-adventures-in-music-organisation/feishin.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/further-adventures-in-music-organisation/feishin.jpg&#34; alt=&#34;A screenshot of Feishin. On the left side of the screen is the album art (Love, Drugs &amp;amp; Misery by Eva Under Fire), with the track, album, artist, year and file format below it. On the right hand side is a tabbed panel, currently showing &amp;#39;Up Next&amp;#39; which shows a queue of music. At the bottom is a standard player interface with play/pause/skip/etc buttons.&#34; loading=&#34;lazy&#34; width=&#34;1723&#34; height=&#34;985&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;Feishin’s ’now playing’ screen&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;What attracted me to it, though, was its built-in support for Navidrome’s smart
playlists. These aren’t even properly exposed in Navidrome’s web UI yet, but
Feishin has a visual editor that lets you create and edit them. I ended up with
an identical system to the playlists I had in Tauon: a “blacklist” playlist
which are tracks I never want to play, a “favourites” playlist, and then an
“everything” playlist which is the entire catalogue minus the blacklist.&lt;/p&gt;
&lt;p&gt;On iOS I’ve settled on &lt;a href=&#34;https://www.reddit.com/r/arpeggiApp/&#34;&gt;Arpeggi&lt;/a&gt;. It’s
not actually on the App Store yet, but available via TestFlight. It’s one of
those really nicely polished apps that are obviously a labour of love, not just
out to do the minimum possible to get your money. It can cache songs offline,
and automatically download entire playlists, and supports all of the standard
Subsonic features like rating, reporting plays back to the server, and so on.&lt;/p&gt;
&lt;p&gt;One unexpected benefit of using a central server is that it handles reporting
plays to &lt;a href=&#34;https://last.fm/&#34;&gt;Last.fm&lt;/a&gt; and &lt;a href=&#34;https://listenbrainz.org/&#34;&gt;ListenBrainz&lt;/a&gt;
instead of the clients. I don’t think I’ve ever bothered to configure my mobile
clients to do that before, and now it just works automagically. I’m relying on
those services more for recommendations and discovery, as that’s something you
lose out on when self-hosting.&lt;/p&gt;
&lt;h3 id=&#34;making-bad-decisions-about-bitrates&#34;&gt;Making bad decisions about bitrates&lt;/h3&gt;
&lt;p&gt;With all the moving, copying, and organising, I started thinking about how
large the collection was. The songs were all in different formats, depending
on when and where I’d picked them up, and I couldn’t tell the difference
between them. I did a test and transcoded an MP3 file down to 128kbps, and
still couldn’t tell the difference. So armed with that sample size of 1, I
transcoded the entire library and saved &lt;em&gt;so much&lt;/em&gt; space.&lt;/p&gt;
&lt;p&gt;I don’t remember which track I did that first test with, but it must have
been a very unlucky pick. I quickly started to notice the distortions caused
by the low bitrate. It was particularly bad in any song with a lot of treble,
and started to really annoy me. I started a painful process of reimporting
things in their original format. Beets came in clutch again, both with the
import process (if you import a duplicate album, it asks if you want to
replace the original, and shows the format and bitrate of them for comparison),
and keeping track of what was left to fix (&lt;code&gt;beet ls bitrate:..128000&lt;/code&gt;).&lt;/p&gt;
&lt;p&gt;Having started paying more attention to the sound quality, I think I’ve been
nibbled on by the audiophile bug. I’ve been trying to get new music in FLAC
format where possible. I’m pretty sure I &lt;em&gt;can’t&lt;/em&gt; tell the difference between
a decent bitrate MP3 and a FLAC, but maybe if I got a better DAC and some nice
headphones…? Someone please hide my credit card!&lt;/p&gt;
&lt;p&gt;As for the library taking up a lot of space, I now realise it’s a price worth
paying. I might need to swap out the server to one with more storage at some
point in the future, but it’s generally worth upgrading every few years with
rented dedicated servers anyway, as you often get more for the same amount of
money. Next upgrade I’ll just make sure there’s an appropriately-sized disk
instead of focusing entirely on RAM and CPU.&lt;/p&gt;
&lt;h3 id=&#34;writing-more-code&#34;&gt;Writing more code&lt;/h3&gt;
&lt;p&gt;Naturally, this whole process spawned several side projects. When browsing
albums in Navidrome, I was a bit upset that the album art wasn’t all a
consistent size. I thought about scripting something to crop them consistently,
but then I had a better idea, and &lt;a href=&#34;https://github.com/csmith/jewelcase&#34;&gt;jewelcase&lt;/a&gt;
was born. It takes the album art, crops it down to a consistent size, and then
renders it inside a jewel case. It applies some slight effects, like adjusting
the colours, rounding the corners, and tweaking the edges so it looks a bit more
real. I don’t know why&lt;sup id=&#34;fnref:4&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:4&#34; role=&#34;doc-noteref&#34;&gt;4&lt;/a&gt;&lt;/sup&gt;, but every time I see the effect it gives me a little
spark of joy.&lt;/p&gt;
&lt;figure class=&#34;image full&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/further-adventures-in-music-organisation/jewelcase.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/further-adventures-in-music-organisation/jewelcase.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/further-adventures-in-music-organisation/jewelcase.jpg&#34; alt=&#34;A comparison screenshot. On the left are four albums, with their unaltered artwork. Some are different aspect ratios. On the right are the same four albums, but the art work is now consistently rendered as though its in a jewel case.&#34; loading=&#34;lazy&#34; width=&#34;1737&#34; height=&#34;702&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;Album art, before and after jewelcase is applied&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The other project is &lt;a href=&#34;https://github.com/csmith/bass&#34;&gt;BASS&lt;/a&gt;, a tool that uses
the Subsonic API to grab information about my music catalogue, and then
generates a “Daily Mix” playlist. It uses a system of weights to select tracks
semi-randomly, biasing towards favourite tracks, those that haven’t been played
much, and a few other criteria. This gives me a nice balance between having my
favourites on repeat, and exploring the full library at random.&lt;/p&gt;
&lt;p&gt;All the weights in BASS are customisable, so I can tweak it to my heart’s
desire, and anyone else can also run it and configure it entirely differently
to me if they want to. I can’t imagine I would’ve been able to do that if I were
using Plex!&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr/&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;I was running a version from git as the stable release had fun dependency
issues on Arch, so it’s possible this wouldn’t be an issue in normal use. I also
started backing up the database, but it’s still annoying to have to restore it. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:1&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:2&#34;&gt;
&lt;p&gt;A “favourites” playlist for me, and a more “family friendly” playlist
for when I’m playing music out loud that has a bit less screaming/swearing/etc. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:2&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:3&#34;&gt;
&lt;p&gt;One of the reasons for starting this was being annoyed by the periodic
sync to my phone, obviously it makes perfect sense to introduce a new, different
periodic sync. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:3&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:4&#34;&gt;
&lt;p&gt;Nostalgia? A little bit of OCD? Both? &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:4&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</content>
    </entry>
    <entry>
        <title>Simple backups with Restic and Hetzner Cloud</title>
        <link href="https://chameth.com/simple-backups-restic-hetzner/"/>
        <updated>2024-12-06T00:00:00Z</updated>
        <id>https://chameth.com/simple-backups-restic-hetzner/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/simple-backups-restic-hetzner/restic.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/simple-backups-restic-hetzner/restic.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/simple-backups-restic-hetzner/restic.png&#34; alt=&#34;The Restic logo — a gopher with two umbrellas.&#34; loading=&#34;lazy&#34; width=&#34;400&#34; height=&#34;400&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;Restic’s mascot, who’s dual-wielding umbrellas to save you from a rainy day.&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I have a confession: for the past few years I’ve not been backing up any of my
computers. Everyone knows that you &lt;em&gt;should&lt;/em&gt; do backups, but actually getting
around to doing it is another story.&lt;/p&gt;
&lt;p&gt;Don’t get me wrong: most of my important things are “backed up” by virtue of
being committed to remote git repositories, or attached to e-mails, or
re-obtainable from the original source, and so on. I don’t think any machine
failing completely would be a disaster for me, but it would certainly be a pain.&lt;/p&gt;
&lt;p&gt;This week I finally got around to doing something, and it ended up being a lot
more straight forward than my previous forays into backup-land.&lt;/p&gt;
&lt;h3 id=&#34;restic&#34;&gt;Restic&lt;/h3&gt;
&lt;p&gt;After soliciting a few opinions, the choice of backup software came down to
either &lt;a href=&#34;https://www.borgbackup.org/&#34;&gt;Borg&lt;/a&gt; or &lt;a href=&#34;https://restic.net/&#34;&gt;Restic&lt;/a&gt;.
I’m pretty sure either would have done what I want, but I leaned towards Restic
for a few reasons: it has a more informative website, it’s written in Go
rather than Python&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:1&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;, and Borg seems to be transitioning between major
releases at the moment&lt;sup id=&#34;fnref:2&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:2&#34; role=&#34;doc-noteref&#34;&gt;2&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;The way Restic works is pretty simple: you initialise a ‘repository’, and can
then call &lt;code&gt;restic backup /some/path&lt;/code&gt; and it’ll get backed up to the repository.
Restic handles keeping different backups separate, and only sending data that’s
changed, and deduplicating, and so on. You basically point it at a thing you
don’t want to lose, and it sorts it out for you. Perfect.&lt;/p&gt;
&lt;p&gt;There are equally straight-forward commands for removing old snapshots
(&lt;code&gt;restic forget&lt;/code&gt;) and verifying backups (&lt;code&gt;restic check&lt;/code&gt;). One of the nice things
about modern backup solutions is they support a whole range of backends. I was
originally going to spin up a small VPS to host my backups, but noticed that
Restic supported S3-compatible stores…&lt;/p&gt;
&lt;h3 id=&#34;hetzner-cloud-object-storage&#34;&gt;Hetzner Cloud Object Storage&lt;/h3&gt;
&lt;p&gt;I host my servers with Hetzner, and was going to use them to spin up a VPS as
well. Despite the “Cloud” branding on a bunch of products, they offer reasonable
prices and good service. A couple of months ago, they started offering
&lt;a href=&#34;https://docs.hetzner.com/storage/object-storage/overview&#34;&gt;S3-compatible object storage&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The pricing isn’t totally straight forward, but for continuous use the “free
quota” amounts to 1TB of storage and 1TB of egress a month. Ingress is free,
as is traffic within the &lt;code&gt;eu-central&lt;/code&gt; region (where all my servers are). That
quota is only awarded when you pay the “base price”, though, which is €4.99 a
month. So basically it’s €5 a month for 1TB of storage and enough egress to
fully restore every single byte. That’s better value than any VPS I can find,
much cheaper than Amazon S3, and about the same as
&lt;a href=&#34;https://www.backblaze.com/cloud-storage&#34;&gt;Backblaze B2&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&#34;getting-them-to-work-together&#34;&gt;Getting them to work together&lt;/h3&gt;
&lt;p&gt;Now you’d think making the tool that supports S3-compatible object storage
work with your S3-compatible object storage would be easy, right? Well not
quite. The library Restic uses to deal with S3 backends has some
strange logic for figuring out the bucket name given a URL. It doesn’t quite
seem to work right, though…&lt;/p&gt;
&lt;p&gt;Hetzner buckets have URLs like &lt;code&gt;s3://bucketname.hel1.your-objectstorage.com&lt;/code&gt;&lt;sup id=&#34;fnref:3&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:3&#34; role=&#34;doc-noteref&#34;&gt;3&lt;/a&gt;&lt;/sup&gt;,
but just passing that to Restic gives an error that the bucket is not specified.
The docs mention there’s an advanced option to make it use the virtual host for
the bucket name: &lt;code&gt;-o s3.bucket-lookup=dns&lt;/code&gt;. But that… also doesn’t work.
What I ended up doing was specifying the URL as
&lt;code&gt;s3://hel1.your-objectstorage.com/bucketname&lt;/code&gt;, and also passing in the &lt;code&gt;dns&lt;/code&gt;
option. The library then seems to muddle its way back to a real, working URL.
I’m not sure why it works like this: maybe it’s just a weird aspect of S3 that
I’m oblivious to?&lt;/p&gt;
&lt;p&gt;The next fun part is that you can configure Restic entirely by using environment
variables, except for that &lt;code&gt;-o s3.bucket-lookup=dns&lt;/code&gt; argument. That has to go
on the command line. I ended up making a little wrapper script to invoke Restic
correctly:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-cp&#34;&gt;#!/bin/sh
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# The password restic will use to encrypt your data. You should generate&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# something nice and secure.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;export&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;RESTIC_PASSWORD&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;repo-password
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# The path to the repository where restic will save the backup. In our case&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# this takes the form `s3:&amp;lt;endpoint&amp;gt;/&amp;lt;bucket&amp;gt;`. Your endpoint might be different&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# to mine depending on the region your bucket is in.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;export&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;RESTIC_REPOSITORY&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;s3:hel1.your-objectstorage.com/bucket-name
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# The access and secret key generated in the &amp;#39;S3 credentials&amp;#39; section of the&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# Hetzner Cloud Console&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;export&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;AWS_ACCESS_KEY_ID&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;hetzner-access-key-id
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;export&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;AWS_SECRET_ACCESS_KEY&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;hetzner-access-key
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# Pass any arguments on to the restic command, along with the magic&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# `s3.bucket-lookup` option we need to resolve the S3 URL properly.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;exec&lt;/span&gt; restic -o s3.bucket-lookup&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;dns &lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;&lt;/span&gt;&lt;span class=&#34;chroma-nv&#34;&gt;$@&lt;/span&gt;&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Then in the backup script I just alias &lt;code&gt;restic&lt;/code&gt; to use the script:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-cp&#34;&gt;#!/bin/bash
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;set&lt;/span&gt; -eu
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# This means any time we use `restic` below, we&amp;#39;ll actually execute our special&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# script which supplies all the env vars and arguments needed to find the&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# repository. Make sure the path matches where you saved the script!&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;alias&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;restic&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;chroma-s1&#34;&gt;&amp;#39;~/.bin/restic&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# Actually do the backup. Each directory in the list below will be backed up&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# separately.&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nv&#34;&gt;dirs&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=(&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;	&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;/some/path/to/backup/&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;	&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;/some/other/path/&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-o&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;for&lt;/span&gt; i in &lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;&lt;/span&gt;&lt;span class=&#34;chroma-si&#34;&gt;${&lt;/span&gt;&lt;span class=&#34;chroma-nv&#34;&gt;dirs&lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;[@]&lt;/span&gt;&lt;span class=&#34;chroma-si&#34;&gt;}&lt;/span&gt;&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;&lt;/span&gt; 
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;do&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;	&lt;span class=&#34;chroma-nb&#34;&gt;echo&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;$i&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;	&lt;span class=&#34;chroma-o&#34;&gt;(&lt;/span&gt;&lt;span class=&#34;chroma-nb&#34;&gt;cd&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;$i&lt;/span&gt; &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; restic --verbose backup .&lt;span class=&#34;chroma-o&#34;&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;done&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# Prune our snapshots. You can tweak the numbers here. Run with `--dry-run` to&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# see what effect any changes would have before actually committing to them. &lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;restic forget --keep-daily &lt;span class=&#34;chroma-m&#34;&gt;7&lt;/span&gt; --keep-weekly &lt;span class=&#34;chroma-m&#34;&gt;10&lt;/span&gt; --keep-monthly &lt;span class=&#34;chroma-m&#34;&gt;24&lt;/span&gt; --keep-yearly &lt;span class=&#34;chroma-m&#34;&gt;10&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The only other notable thing here is that the script changes into the directory
to be backed up. If you give Restic an absolute path, it will create a new
snapshot if the metadata of any folder in the path changes, which is not what
I want.&lt;/p&gt;
&lt;p&gt;If it wasn’t for the S3 URL issues, the whole thing would’ve probably taken
me about half an hour. That’s including setting up the object storage,
installing Restic, and so on. It’s painfully easy. Why didn’t I do this three
years ago?!&lt;/p&gt;
&lt;h3 id=&#34;addendum-a-step-by-step-guide&#34;&gt;Addendum: a step-by-step guide&lt;/h3&gt;
&lt;aside class=&#34;update raised-box&#34;&gt;
  &lt;h5 class=&#34;plain-header&#34;&gt;Update 2025-03-01:&lt;/h5&gt;
  &lt;p&gt;This section was added after the original article was published, following some
helpful feedback. Let me know if you have any problem with these instructions!&lt;/p&gt;
&lt;/aside&gt;
&lt;p&gt;If you want to do this yourself, here’s a quick step-by-step guide:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Log in to the &lt;a href=&#34;https://console.hetzner.cloud&#34;&gt;Hetzner Cloud Console&lt;/a&gt;, and
create a project.&lt;/li&gt;
&lt;li&gt;On the “Object Storage” tab, create a new bucket. Note the name and the
endpoint.&lt;/li&gt;
&lt;li&gt;On the “Security” tab, go to “S3 Credentials” and generate new credentials.
Note down the access key and the secret key.&lt;/li&gt;
&lt;li&gt;Copy the first shell script above, fill in the password (you can pick!),
endpoint, bucket name, access key and secret key.&lt;/li&gt;
&lt;li&gt;Run the script with the &lt;code&gt;init&lt;/code&gt; argument (e.g. &lt;code&gt;~/.bin/restic init&lt;/code&gt;). This
will create a new repository, and only needs to be done once even if you
backup multiple machines.&lt;/li&gt;
&lt;li&gt;Copy the second shell script above, making sure the &lt;code&gt;restic&lt;/code&gt; alias points
at the script you saved in step 4. Change the list of directories to
whatever you want to backup.&lt;/li&gt;
&lt;li&gt;Schedule the script to be run automatically, using crontab or systemd timers
or however you prefer.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;To check everything is working, you can use the &lt;code&gt;snapshots&lt;/code&gt; subcommand to see
a list of saved snapshots. You might also want to try to restore a snapshot
using the &lt;code&gt;restore&lt;/code&gt; subcommand.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr/&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;I’m not trying to be a language snob, but given the choice between two
otherwise equal projects one in Go and one in Python, I’ll take the Go one.
I know I’m not going to have weird library issues down the line, and I’m much
more comfortable rummaging around the source. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:1&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:2&#34;&gt;
&lt;p&gt;With a “Don’t use this in production!” notice on the shiny new version. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:2&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:3&#34;&gt;
&lt;p&gt;Aside: I really hate URLs like that. What is with the trend for completely
generic domains divorced from the service they’re a part of? &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:3&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</content>
    </entry>
    <entry>
        <title>Reproducible Builds and Docker Images</title>
        <link href="https://chameth.com/reproducible-builds-docker-images/"/>
        <updated>2022-02-18T00:00:00Z</updated>
        <id>https://chameth.com/reproducible-builds-docker-images/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image left&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/reproducible-builds-docker-images/dependency.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/reproducible-builds-docker-images/dependency.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/reproducible-builds-docker-images/dependency.png&#34; alt=&#34;Comic showing all modern digital infrastructure is built upon one project by a random person in Nebraska&#34; loading=&#34;lazy&#34; width=&#34;385&#34; height=&#34;489&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;XKCD 2347: Dependency&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;a href=&#34;https://reproducible-builds.org/&#34;&gt;Reproducible builds&lt;/a&gt; are builds which you are able to reproduce byte-for-byte,
given the same source input. Your initial reaction to that statement might be “Aren’t nearly all builds
‘reproducible builds’, then? If I give my compiler a source file it will always give me the same binary, won’t it?”
It &lt;em&gt;sounds&lt;/em&gt; simple, like it’s something that should just be fundamentally true unless we go out of our way to break it,
but in reality it’s actually quite a challenge. A group of Debian developers have been working on reproducible packages
for the best part of a decade and while they’ve made fantastic progress,
&lt;a href=&#34;https://isdebianreproducibleyet.com/&#34;&gt;Debian still isn’t reproducible&lt;/a&gt;. Before we talk about why it’s a hard problem,
let’s take a minute to ponder why it’s worth that much effort.&lt;/p&gt;
&lt;h3 id=&#34;on-supply-chain-attacks&#34;&gt;On supply chain attacks&lt;/h3&gt;
&lt;p&gt;Suppose you want to run some open-source software. One of the many benefits of open-source software is that anyone
can look at the source and, in theory, spot bugs or malicious code. Some projects even have sponsored audits or
penetration tests to affirm that the software is safe. But how do you actually deploy that software? You’re probably
not building from source - more likely you’re using a package manager to install a pre-built version, or downloading
a binary archive, or running a docker image. How do you know whoever prepared those binary artifacts did so from
an un-doctored copy of the source? How do you know a
&lt;a href=&#34;https://en.wikipedia.org/wiki/SourceForge#Controversies&#34;&gt;middle-man hasn’t decided to add malware to the binaries to make money&lt;/a&gt;?&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;Even worse: if the software you’re trying to use includes any dependencies, you have the same issue of trust
with them. Maybe &lt;em&gt;your&lt;/em&gt; supplier isn’t compromising the software, but that doesn’t mean &lt;em&gt;their&lt;/em&gt; supplier isn’t. The
beauty-cum-horror of a supply chain attack is that it can target the weakest link anywhere along the supply chain.
Even if there aren’t any binary files involved, dependencies can still be attacked: what if &lt;code&gt;npmjs.com&lt;/code&gt; or
&lt;code&gt;proxy.golang.org&lt;/code&gt; or &lt;code&gt;github.com&lt;/code&gt; return a different version of a dependency-of-a-dependency when the request
comes from your IP address? It doesn’t even need to be a modified dependency, it could be a perfectly un-tampered,
properly signed copy of the source, just from an older version with a known vulnerability.&lt;/p&gt;
&lt;p&gt;Enter stage left: reproducible builds, here to save the day! If the build process is reproducible then you - or anyone
else on the internet - can perform the same build on the same source and validate the output has the same checksum or
hash. If Debian publish a binary package and an independent re-builder comes up with the exact same build artifact,
there’s a reasonably good chance that the build is good. An attacker would have to compromise both the build machine
and the re-build machine to do anything nefarious. The more re-builders there are, the less feasible a supply chain
attack is.&lt;/p&gt;
&lt;h3 id=&#34;so-why-isnt-software-just-reproducible&#34;&gt;So why isn’t software just reproducible?&lt;/h3&gt;
&lt;h4 id=&#34;compilers&#34;&gt;Compilers&lt;/h4&gt;
&lt;p&gt;As a bit of an experiment, I asked some friends to run the following for me and report the answer:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-nb&#34;&gt;echo&lt;/span&gt; -e &lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;#include &amp;lt;stdio.h&amp;gt;\nint main() { printf(\&amp;#34;Hello\&amp;#34;); return 0; }&amp;#34;&lt;/span&gt; &lt;span class=&#34;chroma-p&#34;&gt;|&lt;/span&gt; &lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  gcc -x c -o hello.out - &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  sha256sum hello.out
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;This compiles a super-simple hello world program and then prints the SHA-256 hash of the resulting binary. Here are
the results:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hash&lt;/th&gt;
&lt;th&gt;System&lt;/th&gt;
&lt;th&gt;GCC&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1f62feab5a06861dc575201d807781926d1ae49fb113da018fde8b670a1346f7&lt;/td&gt;
&lt;td&gt;Arch&lt;/td&gt;
&lt;td&gt;11.2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;b8e6f2c7082be69f65ffa5e7a3d749eb47866a1b2e1ec19efb63cc59a8b160cd&lt;/td&gt;
&lt;td&gt;Debian&lt;/td&gt;
&lt;td&gt;8.3.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cbad2e47a22c234b5e7fa55e029a8db4d64ac7a962e2176bd2e1373d78954088&lt;/td&gt;
&lt;td&gt;Debian&lt;/td&gt;
&lt;td&gt;8.3.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;e0f6bbc13b29fea8cfa2a975ba4661e781323298aec166c8311d342e6f93c4a6&lt;/td&gt;
&lt;td&gt;Alpine&lt;/td&gt;
&lt;td&gt;10.3.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;e379156895e06c7a0bf18ac4d648860edcb2655576b0ab9fab172bd6c8b92075&lt;/td&gt;
&lt;td&gt;Debian&lt;/td&gt;
&lt;td&gt;10.2.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7ffdaee4eb64e016b89dc5e54d2c8eebab3cebafe2c7aa97de627b5972ecea46&lt;/td&gt;
&lt;td&gt;Debian&lt;/td&gt;
&lt;td&gt;11.2.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8ae52cc166743b6ae1eb3e14179ef33de5061a04237f8f97088c896c41a2f698&lt;/td&gt;
&lt;td&gt;Arch&lt;/td&gt;
&lt;td&gt;11.1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8ae52cc166743b6ae1eb3e14179ef33de5061a04237f8f97088c896c41a2f698&lt;/td&gt;
&lt;td&gt;Arch&lt;/td&gt;
&lt;td&gt;11.1.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;As you can see, there are barely any duplicates. Even the same version of GCC on the same OS sometimes produces
different results. And this is the most basic program I could write! Differences arise from the compiler version,
the build flags, the libraries installed, and a whole host of other factors. If you compile a Go application instead of
a C one, then by default the compiler will include debug information in the binary. This includes the full path to the
source file on disk, so building a project in &lt;code&gt;/home/chris/&lt;/code&gt; will produce a different binary to building the same
source in &lt;code&gt;/tmp&lt;/code&gt;. Future versions of Go are also going to stamp in other meta-data such as VCS info, so building inside
and outside a Git repository will produce different binaries.&lt;/p&gt;
&lt;h4 id=&#34;archives&#34;&gt;Archives&lt;/h4&gt;
&lt;p&gt;Compilers are only half the problem. Build processes are usually multistep, involving compiling, moving, compressing,
and so on. Consider creating an archive of a file:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;repeat &lt;span class=&#34;chroma-m&#34;&gt;4&lt;/span&gt; touch hello &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; tar zcf hello.tgz hello &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; sha256sum hello.tgz &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; sleep 0.5
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;f3d5c56f6b8089de95d62d060e6ffcbbad26875807ae7bc253f07cd097ea61be  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;ab67f2e865b5afa87d9b2434d92b0c271b3cf730fa85988f84852551749ba6ed  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;ab67f2e865b5afa87d9b2434d92b0c271b3cf730fa85988f84852551749ba6ed  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;738678c9650b10fd83636997dd1aba4016bbf0ec5ebf3dfd4ef75d770b56e23b  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Any file added to a tar takes with it a timestamp, so the build is only reproducible if it happens at the exact same
time! We can make this reproducible by forcing &lt;code&gt;tar&lt;/code&gt; (and the same goes for &lt;code&gt;zip&lt;/code&gt; and most other archive formats) to
set a certain timestamp on the files:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;repeat &lt;span class=&#34;chroma-m&#34;&gt;4&lt;/span&gt; touch hello &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; tar --mtime 2022-02-18T01:00 -zcf hello.tgz hello &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; sha256sum hello.tgz &lt;span class=&#34;chroma-o&#34;&gt;&amp;amp;&amp;amp;&lt;/span&gt; sleep 0.5 
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;081060a900beff2a6aad9957a8cbb8792f8db7904f86b318dbf26b682a2d3f0a  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;081060a900beff2a6aad9957a8cbb8792f8db7904f86b318dbf26b682a2d3f0a  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;081060a900beff2a6aad9957a8cbb8792f8db7904f86b318dbf26b682a2d3f0a  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;081060a900beff2a6aad9957a8cbb8792f8db7904f86b318dbf26b682a2d3f0a  hello.tgz
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;In a real build there are basically two approaches here: you can set it to a pre-defined value (like the unix epoch),
or you can set it to match the modification time of the source files. The former is easiest, but the latter is more
cosmetically and semantically appealing.&lt;/p&gt;
&lt;h4 id=&#34;iteration-order&#34;&gt;Iteration order&lt;/h4&gt;
&lt;p&gt;So we’ve pinned our build environment, we’re manipulating timestamps when adding files to archives, now what? Imagine
part of the build process involves looping through all the files in a directory and doing &lt;em&gt;something&lt;/em&gt;. What order do
these files get iterated in? Well, that very much depends on the filesystem and perhaps when the files themselves
were created. To ensure this is reproducible we need to explicitly sort any such operation so that it’s always
consistent. This iteration could be happening in a tool that’s called by another tool that’s called by a build script,
so the fix isn’t necessarily straight-forward.&lt;/p&gt;
&lt;aside class=&#34;sidenote raised-box&#34;&gt;
  &lt;h5 class=&#34;plain-header&#34;&gt;Side note: a bug war story&lt;/h5&gt;
  &lt;p&gt;I’ve personally been victim to this kind of non-determinism. I was working on an Android app, and committed a new
test that worked fine on my machine, and worked fine on the CI server. But it failed consistently for a colleague.&lt;/p&gt;
&lt;p&gt;We both did fresh checkouts of the source, and ran the tests. Mine passed, his failed. He sent me an archive of
his checkout in case there was something weird going on there, and the tests passed on my machine. We compared
hashes of our checkouts, and they were the same. It was obviously environmental somehow, but everything else worked
fine, and the build system went to great pains to ensure things were the same.&lt;/p&gt;
&lt;p&gt;After a &lt;em&gt;lot&lt;/em&gt; of debugging, I worked out that his test was running with a different version of a library to me,
despite the libraries being defined in the build files and the build files being identical. After &lt;em&gt;even more&lt;/em&gt;
debugging it turned out there were two versions of the library on the classpath, and the ordering of them was
different between my machine and his.&lt;/p&gt;
&lt;p&gt;The actual issue turned out to be that the build tool generated the classpath by iterating over the library
files, and that iteration was done in order of file creation time. The two libraries were added at different points
in the project history, so the creation time in your local cache depended on which versions of the app you’d built
in the past. With no cache everything worked as expected but there was a slim range of commits where only one
library was in use, and if you had run the tests during that period your cache was effectively poisoned.&lt;/p&gt;
&lt;p&gt;We fixed the issue by excluding the older version of the library (which was being pulled in as a transient dependency),
and filed a bug against the build tool to make the classpath properly deterministic. I think that stands as the most
difficult to diagnose bug I’ve ever dealt with.&lt;/p&gt;
&lt;/aside&gt;
&lt;p&gt;Interestingly, if you iterate over a map in Go, the iteration is &lt;em&gt;deliberately&lt;/em&gt; non-deterministic. That’s an attempt
to defeat &lt;a href=&#34;https://www.hyrumslaw.com/&#34;&gt;Hyrum’s Law&lt;/a&gt; and prevent developers from relying on whatever the current
behaviour happens to be. This actually makes it easier to make things reproducible as the problem is loud and
in-your-face, rather than subtle and hard to spot.&lt;/p&gt;
&lt;h4 id=&#34;other-sources&#34;&gt;Other sources&lt;/h4&gt;
&lt;p&gt;There’s an awful lot of other places that non-determinism can come from. If the app pulls in dependencies, their
versions have to be pinned, otherwise your build changes depending on the latest release of that dependency. If
the build process pulls any information from a website, it’s liable to change. Hopefully the website is under your
control so that you can version the resource and pin that version. Obviously, anything to do with dates or the
current user will probably cause problems. Timezones and locales can cause subtle differences.&lt;/p&gt;
&lt;h3 id=&#34;what-about-docker&#34;&gt;What about Docker?&lt;/h3&gt;
&lt;p&gt;Docker comes with some good and some bad points for reproducibility. The biggest advantage is that it inherently
completely describes the build environment; it should work exactly the same from one system to another, even across
different OS families. The biggest drawback is it sprays timestamps around like no-one’s business. Each layer in
a container image is a &lt;code&gt;.tar.gz&lt;/code&gt; file, meaning each file within it is timestamped as discussed above. Making an image
involves a lot of copying of files around, so these timestamps invariably end up causing reproducibility issues.&lt;/p&gt;
&lt;p&gt;Even worse than timestamps in the filesystem, the image format also contains some meta-data that includes the
timestamp at which each layer was built. That means even if you go out of your way to set the timestamp of every
single file in your image, the image itself will be different every time you rebuild it. There is no way to deal
with this in Docker, which is a very sad state of affairs. Fortunately, &lt;a href=&#34;https://buildah.io/&#34;&gt;Buildah&lt;/a&gt; provides
a &lt;code&gt;--timestamp&lt;/code&gt; flag for &lt;em&gt;its&lt;/em&gt; build commands; this not only sets the layer timestamp but also the creation
timestamp of any file within the layer.&lt;/p&gt;
&lt;p&gt;The other major issue that affects Docker images is the pinning of packages pulled in by package managers. An awful
lot of images are based on Alpine or Debian derivatives, and use &lt;code&gt;apk&lt;/code&gt; or &lt;code&gt;apt&lt;/code&gt; to install dependencies. These need
to have a version specified as otherwise the package manager will just pull in the latest at the time of the build.
But this isn’t quite enough: you also need to pin the version of any packages that they depend on, recursively.
This means flattening the entire package hierarchy and installing all the packages explicitly and with pinned
versions.&lt;/p&gt;
&lt;p&gt;One more wrinkle in the package management space is that Alpine don’t keep old packages in their main repositories.
If you have a Docker image with pinned alpine packages in, it will stop building if the package is updated. This
isn’t necessarily fatal to making a reproducible build — as long as it’s reproducible for its useful lifetime,
I don’t really see an issue.&lt;/p&gt;
&lt;p&gt;Honestly, though, the biggest issue with making Docker images reproducible is getting people to care. Dockerfiles
are a relatively new way of packaging software, and there’s no centralised organisation like you find with Linux
distributions. There are enough challenges that most casual packagers aren’t going to bother, and no real
incentive for them to. That won’t stop me trying, though!&lt;/p&gt;
</content>
    </entry>
    <entry>
        <title>Artisanal Docker images</title>
        <link href="https://chameth.com/artisanal-docker-images/"/>
        <updated>2022-02-05T00:00:00Z</updated>
        <id>https://chameth.com/artisanal-docker-images/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/artisanal-docker-images/artisanal-containers.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/artisanal-docker-images/artisanal-containers.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/artisanal-docker-images/artisanal-containers.jpg&#34; alt=&#34;Shelf showing a variety of artisanal containers&#34; loading=&#34;lazy&#34; width=&#34;300&#34; height=&#34;432&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;Artisanal containers…&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I run a fair number of services as docker containers. Recently, I’ve been moving away from pre-built images
pulled from Docker Hub in favour of those I’ve hand-crafted myself. If you’re thinking “that sounds like a
lot of effort”, you’re right. It also comes with a number of advantages, though, and has been a fairly fun
journey.&lt;/p&gt;
&lt;h3 id=&#34;the-problems-with-docker-hub-and-its-images&#34;&gt;The problems with Docker Hub and its images&lt;/h3&gt;
&lt;h4 id=&#34;rate-limits&#34;&gt;Rate limits&lt;/h4&gt;
&lt;p&gt;For the last few years, I’ve been getting increasingly unhappy with Docker Hub itself. Docker-the-technology
is wonderful, but Docker-the-company has been making some rather large missteps. The biggest and most impactful
of these has been introducing “pull rate” limits. At the time of writing, if you want to just pull a public image
without logging in then you are limited to 100 pulls every 6 hours. If you log in then you’re limited to 200 pulls
per 6 hours, but it’s account wide. This might seem like a big enough number, but I repeatedly hit it and there
is no way to actually audit what is causing it. I have various containers that may all pull images at arbitrary
times (e.g. continuous integration build agents), and the only information you get back from Docker Hub is the
number of pulls remaining.&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;Obviously, I could start paying Docker Hub for a “Pro” plan. That gets you 5,000 pulls per day for $7/month.
The downside is that every docker client would have to be authenticated, which presents a fair annoyance in
terms of credential management. I also don’t really like how they positioned the service as a public utility
with special treatment in the docker software, and then start tightening the ratchet to make money.&lt;/p&gt;
&lt;h4 id=&#34;bad-images&#34;&gt;“Bad” images&lt;/h4&gt;
&lt;p&gt;I’m fairly opinionated about what a container image should look like: most importantly it should run just a
single process, and only include the bare minimum dependencies required for that. Other people think differently,
and it’s very hard to tell at a glance whether an image on Docker Hub contains just the application you want,
or whether it also bundles MySQL, Redis, Elasticsearch, and a partridge in a pear tree. Some people want that
kind of thing, but I really don’t. It’s also very hard to tell whether an image is officially endorsed by the
upstream project, and where the source Dockerfile is. This used to be better because most projects used Docker Hub’s
automatic builds, but they’re now a “pro” feature.&lt;/p&gt;
&lt;p&gt;I quite often found that I’d be looking for an image for X, and there would be 5-10 images from different users.
None of them looked official, some of them were out-of-date, some bundled the kitchen sink. Even when one looked
good, it’s a bit of a gamble whether the author is going to keep it updated or not.&lt;/p&gt;
&lt;h4 id=&#34;doijanky&#34;&gt;Doijanky&lt;/h4&gt;
&lt;p&gt;The rate limits and other problems were annoying, but they weren’t really annoying enough to force me to do
anything about it. The straw that broke the camel’s back came later: I was looking at the
&lt;a href=&#34;https://hub.docker.com/_/golang&#34;&gt;official golang images&lt;/a&gt;, and noticed that all the tags were pushed by a
random user account called “doijanky”:&lt;/p&gt;
&lt;figure class=&#34;image center&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/artisanal-docker-images/doijanky.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/artisanal-docker-images/doijanky.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/artisanal-docker-images/doijanky.png&#34; alt=&#34;An &amp;#39;official&amp;#39; Docker Hub image pushed by user &amp;#39;doijanky&amp;#39;&#34; loading=&#34;lazy&#34; width=&#34;786&#34; height=&#34;249&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;An ‘official’ Docker Hub image pushed by user ‘doijanky’&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;I, perhaps naively, assumed that official images were built on Docker Hub’s own infrastructure. Why would
all the Golang images be attributed to this user? Checking out their profile, they’re simply identified as
a “Community User” like everyone else, with no repositories of their own. The only thing in the profile is
their homepage, which is a link to a Jenkins dashboard: &lt;a href=&#34;https://doi-janky.infosiftr.net/&#34;&gt;https://doi-janky.infosiftr.net/&lt;/a&gt;. It appears
legitimate: “Infosiftr” are a container consultancy and the dashboard is linked to from the README in the
official images git repository, but I find it baffling that they’re using third-party infrastructure and
a normal user account (with a dubious name) to push these images. There doesn’t seem to be a good way to
verify what you pull corresponds to the Dockerfile it came from; if infosiftr wanted to inject something
into the build they could happily do so, and who knows how good their infosec posture is? If someone got
access to the “doijanky” account, how long could they upload malicious images before someone noticed?&lt;/p&gt;
&lt;p&gt;This little roller-coaster ride from “are all the official images compromised?!” to “oh, no, they’re not,
it’s all just awful” finally convinced me to look at building my own images from scratch.&lt;/p&gt;
&lt;h3 id=&#34;the-implementation-templating-with-contempt&#34;&gt;The implementation: templating with contempt&lt;/h3&gt;
&lt;p&gt;One of the big issues I needed to tackle was how to deal with updates. I didn’t want to have to go and
edit a file every time some minor release was made of some software, or every time there was a security
vulnerability in a common library. The official images use a shell-scripting based system to check for
updates and generate Dockerfiles, I decided to do something similar but with Go templates. The result is
a tool called &lt;a href=&#34;https://github.com/csmith/contempt&#34;&gt;contempt&lt;/a&gt;. It takes a template like:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;FROM&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-s&#34;&gt;{{image&lt;/span&gt; &lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;golang&amp;#34;&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;}}&lt;/span&gt; AS build&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;ARG&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;TAG&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;{{github_tag &amp;#34;&lt;/span&gt;example/project&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;}}&amp;#34;&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;RUN&lt;/span&gt; apk add --no-cache &lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-o&#34;&gt;{{&lt;/span&gt;range &lt;span class=&#34;chroma-nv&#34;&gt;$key&lt;/span&gt;, &lt;span class=&#34;chroma-nv&#34;&gt;$value&lt;/span&gt; :&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt; alpine_packages &lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;git&amp;#34;&lt;/span&gt; -&lt;span class=&#34;chroma-o&#34;&gt;}}&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-o&#34;&gt;{{&lt;/span&gt;&lt;span class=&#34;chroma-nv&#34;&gt;$key&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;}}={{&lt;/span&gt;&lt;span class=&#34;chroma-nv&#34;&gt;$value&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;}}&lt;/span&gt;&lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-o&#34;&gt;{{&lt;/span&gt;end&lt;span class=&#34;chroma-o&#34;&gt;}}&lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;;&lt;/span&gt; &lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# ...&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Contempt has support for getting information from a variety of sources. In this case, it’s getting
the latest digest of another Docker image, the latest tag from a Git repository, and the latest version
of an alpine package and all its dependencies. The resulting Dockerfile looks something like this:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c&#34;&gt;# Generated from https://github.com/csmith/dockerfiles/blob/master/miniflux/Dockerfile.gotpl&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c&#34;&gt;# BOM: {&amp;#34;apk:brotli-libs&amp;#34;:&amp;#34;1.0.9-r5&amp;#34;,&amp;#34;apk:busybox&amp;#34;:&amp;#34;1.34.1-r4&amp;#34;, &amp;lt;snip&amp;gt; }&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;FROM&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-s&#34;&gt;reg.c5h.io/golang@sha256:ac8fa5f4078b0a697796b5d741&lt;/span&gt;... AS build&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;ARG&lt;/span&gt; &lt;span class=&#34;chroma-nv&#34;&gt;TAG&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;&lt;span class=&#34;chroma-s2&#34;&gt;&amp;#34;2.0.35&amp;#34;&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;RUN&lt;/span&gt; apk add --no-cache &lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        brotli-libs&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;1.0.9-r5&lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-nv&#34;&gt;busybox&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;1.34.1-r4&lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-c1&#34;&gt;# &amp;lt;snip&amp;gt;&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-nv&#34;&gt;pcre2&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;10.39-r0&lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-nv&#34;&gt;zlib&lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;=&lt;/span&gt;1.2.11-r3&lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;        &lt;span class=&#34;chroma-p&#34;&gt;;&lt;/span&gt; &lt;span class=&#34;chroma-se&#34;&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;# ...&lt;/span&gt;&lt;span class=&#34;chroma-err&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;I’ve cut out the longer parts for readability. You can see that it pins all versions of the alpine packages in use,
as well as the base image. This ensures that if you build from the same Dockerfile at a later time it will build the
same image (or will fail entirely, as Alpine doesn’t keep their old packages around indefinitely). It also produces
a “bill of materials” as a really long JSON-encoded comment. If you let contempt commit the Dockerfile it uses the
BOM to generate useful commit messages like: &lt;code&gt;[project] apk:busybox: 1.34.1-r3-&amp;gt;1.34.1-r4&lt;/code&gt;, so you can see exactly
what changed.&lt;/p&gt;
&lt;p&gt;Contempt also has support for building and pushing images whenever it changes the Dockerfile. I use it in a
GitHub action that runs daily to check all my images are up-to-date and push those that aren’t. It understands
the dependencies between images (by pre-analysing the templates) so it will always check and build base images
before ones that require them. This means an update to, say, the “alpine” base image will cause anything that
depends on it to get updated at the same time, ensuring security updates are rolled out promptly.&lt;/p&gt;
&lt;h3 id=&#34;the-result&#34;&gt;The result&lt;/h3&gt;
&lt;p&gt;You can see my collection of lovingly hand-crafted Dockerfiles in my &lt;a href=&#34;https://github.com/csmith/dockerfiles&#34;&gt;dockerfiles&lt;/a&gt;
repository.&lt;/p&gt;
&lt;p&gt;There are a number of advantages to handwriting all the images I use. The obvious one is that they’re all built
how I want: there are no extraneous dependencies, they’re all based on the same small set of base images (rather
than pulling around 10 different versions of debian), nothing tries to also run a DBMS in its container, etc.&lt;/p&gt;
&lt;p&gt;This level of customisation goes further, though. Because I’m packaging the software myself, I can tweak how it’s
built to fit my needs. A couple of things I run need their own TLS certificates separate from my normal HTTPS
setup, so I bake my &lt;a href=&#34;https://github.com/csmith/certwrapper/&#34;&gt;certwrapper&lt;/a&gt; tool in to manage those; I can even set the
build flags on certwrapper to only enable the particular DNS provider I personally need (thus avoiding dragging in
clients for AWS, GCP, etc). Some software like Hashicorp Vault has an optional web interface that I don’t need,
so I simply don’t enable it in the build. These changes save build time, reduce image sizes, in some cases improve
runtime performance, and generally reduce the attack surface of what’s running in the container.&lt;/p&gt;
&lt;p&gt;It’s also been a great way to learn more about how software is distributed. Writing a Dockerfile is not that distant
from writing a PKGBUILD file for an Arch Linux package, or the equivalent for other distributions. In a couple of
instances I’ve googled how to solve a particular issue, and found an Arch or Void linux maintainer asking the upstream
project about the exact same issue.&lt;/p&gt;
&lt;p&gt;Finally, all the images I build I push to my own registry so there are obviously no rate limiting issues.
Standing up a service (assuming the Dockerfile has been written!) is amazingly quick because the base layers are all
shared and cached, and the registry is a lot physically closer than Docker Hub. Bootstrapping this whole thing becomes
an interesting problem because the image for the registry is stored on the registry, but I’ll leave that discussion for
another post…&lt;/p&gt;
</content>
    </entry>
</feed>
