<?xml version="1.0" encoding="utf-8"?>
<?xml-stylesheet href="/feeds.xsl" type="text/xsl"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:base="https://chameth.com/">
    <title>Chameth.com - posts like debugging-beyond-the-debugger, the-ethics-of-llms, why-you-should-be-using-https but not docker-automatic-nginx-proxy, migrating-from-github-to-forgejo</title>
    <subtitle>Personal homepage of Chris Smith</subtitle>
    <link href="https://chameth.com/feeds/posts/like/debugging-beyond-the-debugger,the-ethics-of-llms,why-you-should-be-using-https/unlike/docker-automatic-nginx-proxy,migrating-from-github-to-forgejo/" rel="self"/>
    <link href="https://chameth.com/"/>
    <icon>https://chameth.com/favicon.png</icon>
    <updated>2025-07-16T00:00:00Z</updated>
    <id>https://chameth.com/</id>
    <author>
        <name>Chris Smith</name>
    </author>
    <entry>
        <title>How tech companies failed to build the Star Trek computer</title>
        <link href="https://chameth.com/how-tech-companies-failed-to-build-the-star-trek-computer/"/>
        <updated>2025-07-16T00:00:00Z</updated>
        <id>https://chameth.com/how-tech-companies-failed-to-build-the-star-trek-computer/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/how-tech-companies-failed-to-build-the-star-trek-computer/enterprise-computer-room.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/how-tech-companies-failed-to-build-the-star-trek-computer/enterprise-computer-room.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/how-tech-companies-failed-to-build-the-star-trek-computer/enterprise-computer-room.jpg&#34; alt=&#34;Still from an episode of Star Trek: The Next Generation, with various characters stood around in a computer core room&#34; loading=&#34;lazy&#34; width=&#34;500&#34; height=&#34;376&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;A computer core room on the Enterprise-D&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In most Star Trek series, the ship or station computer is ever-present in the
background, waiting to be called on by the main characters&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:1&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;. It nearly
always does exactly the right thing, and there’s little limit to the functions
it can perform. Take this mundane example from DS9:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;KIRA: Computer, establish link with the Bajoran Medical Index for the Northwestern District. &lt;br/&gt;
COMPUTER: Link established. &lt;br/&gt;
KIRA: Access all information on Doctor Surmak Ren. &lt;br/&gt;
COMPUTER: There are no records matching that name. &lt;br/&gt;
KIRA: Try the Northeastern District, same search. &lt;br/&gt;
COMPUTER: Doctor Surmak Ren, currently serving as Chief Administrator of the Ilvian Medical Complex. &lt;br/&gt;
KIRA: Computer, open a channel to the Ilvian Medical Complex. Administrator’s office.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The computer is doing some kind of networking to a database only identified by
name. It does a search and summarises the lack of results. It then repeats the
process with another database, and succinctly announces the results. Finally,
it opens a communication channel to a specific room in a facility, based only
on its name.&lt;/p&gt;
&lt;p&gt;This whole interaction is remarkably boring&lt;sup id=&#34;fnref:2&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:2&#34; role=&#34;doc-noteref&#34;&gt;2&lt;/a&gt;&lt;/sup&gt;. Kira doesn’t have to know
any URLs or API endpoints, or what protocol she wants to use. She doesn’t have
to open a specific app and then login and then try the query again. She just
says what she wants and the computer does it.&lt;/p&gt;
&lt;p&gt;It seems like this should be one of the most easily obtainable bits of sci-fi
wizardry with our current technology. We have multiple massive companies
throwing lots of money at digital assistants, LLMs that are improving at an
insane rate, but we’re somehow not even close to the usability or usefulness of
the Trek computers. What gives?&lt;/p&gt;
&lt;h3 id=&#34;boring-is-well-boring&#34;&gt;Boring is, well, boring.&lt;/h3&gt;
&lt;p&gt;Larry Page once said something that might help explain it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The Star Trek computer doesn’t seem that interesting. They ask it random
questions, it thinks for a while. I think we can do better than that.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is the same Larry Page that founded Google, whose mission statement is
“to organize the world’s information and make it universally accessible and
useful”. Of all people, surely he should find an omnipresent computer that can
answer ‘random questions’ interesting?! It seems like it should be the epitome
of Google’s mission!&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;Google’s “better than that” seems to have been to stuff LLMs into every product
they can, even when you don’t want them there. Even when they’re worse than the
normal content they displace. These things look &lt;em&gt;exciting&lt;/em&gt; when they’re part of
a scripted demo at Google I/O, but they fall flat and just get in the way when
they’re exposed to the reality of day-to-day use.&lt;/p&gt;
&lt;p&gt;The Star Trek computer is the opposite: it isn’t snazzy, but it is genuinely
useful. That means it’s not an attractive target for the company execs who want
marketing opportunities, and it’s not appealing for engineers who need to
demonstrate “impact”. But even if Google did try to make the Trek computer,
there are other problems…&lt;/p&gt;
&lt;h3 id=&#34;assistants-need-to-be-free&#34;&gt;Assistants need to be free&lt;/h3&gt;
&lt;p&gt;A significant amount of tech companies’ business models currently revolves
around trapping users in walled gardens. They want you using &lt;em&gt;their&lt;/em&gt; ecosystem;
that way they get more data from you, and you’re more likely to spend more money
on their other offerings that work together. There’s barely any incentive to
allow any kind of interoperability with other platforms outside carefully
contracted integrations.&lt;/p&gt;
&lt;p&gt;I remember trying to help a family member move their photos from iCloud to
Google Photos. At one point they turned around and said, exasperated, “why is
this so hard? Aren’t they both in the cloud?!”. It’s easy to dismiss that as
someone who hasn’t quite grasped the fundamental idea that “the cloud” is just
someone else’s computers, but that’s not the whole story. There’s no reason why
there shouldn’t be a quick and easy transfer: both services already allow
uploading and downloading, there’s just no incentive for the companies involved
to make it so&lt;sup id=&#34;fnref:3&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:3&#34; role=&#34;doc-noteref&#34;&gt;3&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;These kinds of misaligned incentives and walled garden business models cause
even more problems when it comes to digital assistants. Siri is basically never
going to be able to interact with, say, your Google Drive; &lt;del&gt;Bard&lt;/del&gt; Gemini
is never going to be able to send a message via iMessage. Even when there are
appropriately blessed interactions, they’re so clunky. Can you imagine Captain
Picard saying “Computer, ask the turbolift skill to take me to deck 5”?&lt;/p&gt;
&lt;h3 id=&#34;someone-elses-computer&#34;&gt;Someone else’s computer&lt;/h3&gt;
&lt;p&gt;Software issues aside, there’s still a key difference between the Star Trek
computers and our current batch of digital assistants: where they run. The
Trek computers are all housed within the ship or station they serve; they can
connect elsewhere to gather information, but they run entirely independently.
If they go wrong, a local engineer can go in and fix things. While some of our
assistants may have physical hardware in your home, they don’t work without
a vast cloud apparatus behind them. If your Internet connection fails, they
become paperweights. If the company running them decide to remove some
functionality you depend on, you have no recourse.&lt;/p&gt;
&lt;p&gt;That kind of helplessness isn’t limited to assistants, either. There’s a rapidly
growing trend of being unable to modify or repair hardware you fully own and
control. Part of this is just that they’re becoming more complex: it’s a lot
harder to replace a microchip than a gear, but companies are also going out
of their way to make it more difficult for users through draconian DRM
regimes&lt;sup id=&#34;fnref:4&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:4&#34; role=&#34;doc-noteref&#34;&gt;4&lt;/a&gt;&lt;/sup&gt; and aggressive intellectual property enforcement. If the US Navy
can’t repair their own equipment because a corporation says so, what hope do
consumers have?&lt;/p&gt;
&lt;p&gt;We’re approaching a point where you don’t actually own anything. Software
is cloud and subscription based, hardware is unrepairable. Even cars can
be remotely updated and have features added or removed. The Federation wouldn’t
allow a third party control over their ships&lt;sup id=&#34;fnref:5&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:5&#34; role=&#34;doc-noteref&#34;&gt;5&lt;/a&gt;&lt;/sup&gt;, so why are we so happy to
put up with it in everything we consume?&lt;/p&gt;
&lt;h3 id=&#34;a-small-ray-of-hope&#34;&gt;A small ray of hope?&lt;/h3&gt;
&lt;p&gt;The most promising way of tackling all of these problems is through legislation.
The EU’s &lt;a href=&#34;https://digital-markets-act.ec.europa.eu/index_en&#34;&gt;Digital Market Act&lt;/a&gt;
is an attempt to force ‘gatekeepers’ like Google, Apple and Meta, to allow
third-party access to their services. It seems like a pretty reasonable
approach, but the tech companies are unsurprisingly resisting it. Apple in
particular have gone out of their way to refuse to comply, and when forced to
do so have limited the functionality to people in Europe.
Still, the DMA is a promising start, and if similar legislation is introduced
(and robustly enforced) elsewhere it might start forcing companies to behave a
bit better.&lt;/p&gt;
&lt;p&gt;There are also smaller companies that actually do the right thing.
&lt;a href=&#34;https://frame.work/gb/en&#34;&gt;Framework&lt;/a&gt; make laptops that are user-serviceable;
&lt;a href=&#34;https://www.fairphone.com/&#34;&gt;Fairphone&lt;/a&gt; do the same for mobile phones. Smaller
software companies provide useful, open APIs. The average person on the street
will probably have never heard of these, unfortunately, but they do still
exist. Maybe as the bigger tech companies tighten the screws more, people will
turn to alternatives like this? Or maybe we’ll just keep accepting that our
computers work for everyone but us?&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr/&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;Unless, of course, the computer is playing the role of the episode’s
MacGuffin and has contracted space-computer-COVID or something, then it’s a lot
less in-the-background. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:1&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:2&#34;&gt;
&lt;p&gt;It’s almost like it only exists to move the plot along. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:2&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:3&#34;&gt;
&lt;p&gt;You can generally export your data, thanks to a combination of legislation
and efforts like Google’s “Data Liberation Front”, but I’ve never seen an export
format that could then just be imported into an equivalent commercial product. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:3&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:4&#34;&gt;
&lt;p&gt;Oh, you’ve changed the screen on your iPhone? Better hope it can do the
secret handshake with the Apple hardware. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:4&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:5&#34;&gt;
&lt;p&gt;I think there might actually have been an episode where that did in fact
happen. We’ll just ignore that as a plot contrivance. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:5&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</content>
    </entry>
    <entry>
        <title>The Ethics of LLMs</title>
        <link href="https://chameth.com/the-ethics-of-llms/"/>
        <updated>2025-06-22T00:00:00Z</updated>
        <id>https://chameth.com/the-ethics-of-llms/</id>
        <content xml:lang="en" type="html">&lt;p&gt;I’ve written about LLMs a few times recently, carefully dodging the issue of
ethics each time. I didn’t want to bog down the other posts
with it, and I wanted some time to think over the issues. Now I’ve had time
to think, it’s time to remove my head from the sand. There are a lot of
different angles to consider, and a lot of it is more nuanced than is often
presented. It’s not all doom and gloom, and it’s also not the most amazing
thing since sliced bread. Who would have thought?&lt;/p&gt;
&lt;p&gt;It’s worth noting that I’m just setting out my position here. I’m not trying
to convince you to change your mind. To set the scene a bit, I mainly use
Claude Code as a programming tool. I use the Claude chat interface sometimes to
proofread things, do random one-off data analysis, or help organise things.
More rarely I’ll try to use it to brainstorm things, or recommend things,
or do more “creative” things, but I don’t trust it enough in those domains
to do it often. I don’t use it for research or as a Google replacement&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:1&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;,
which I recognise probably makes me a weird half-in-the-water, half-out-of-it
class of user.&lt;/p&gt;
&lt;h3 id=&#34;copyright--corporate-control&#34;&gt;Copyright &amp;amp; Corporate Control&lt;/h3&gt;
&lt;p&gt;One of the key issues, and something that is being prosecuted in several court
cases right now, is how LLMs interact with the copyright system. And by
“interact with” I mean “run roughshod all over”. It seems pretty obvious from
my lay perspective that if having 10 seconds of pop music in the background of
a YouTube video is copyright infringement, then
&lt;a href=&#34;https://www.theguardian.com/technology/2025/jan/10/mark-zuckerberg-meta-books-ai-models-sarah-silverman&#34;&gt;Meta pirating books via BitTorrent&lt;/a&gt;
must also be.&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;That said, the entire copyright system as it exists now is fundamentally
broken. I don’t mind the idea of individual creators having protection, but I absolutely hate
how copyright is wielded as a blunt instrument by massive corporations to
intimidate individuals and to try to earn more money. If a law enables a
huge conglomerate to sue the author of an open source project for a
&lt;a href=&#34;https://en.wikipedia.org/wiki/Arista_Records_LLC_v._Lime_Group_LLC&#34;&gt;theoretical 75 trillion dollars&lt;/a&gt;
then that is a bad law. Copyright today has been twisted well beyond what it
was meant to be originally, thanks to lobbying by people and businesses with
lots of money that want to keep it.&lt;/p&gt;
&lt;p&gt;So if copyright law is not fit for purpose, are LLMs fine? Well, not quite.
The argument that LLM training is a bit like a child learning to read holds
some merit, but it falls flat when LLMs can regurgitate the input material
verbatim&lt;sup id=&#34;fnref:2&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:2&#34; role=&#34;doc-noteref&#34;&gt;2&lt;/a&gt;&lt;/sup&gt;. Personally I don’t care if they steal all the content of the
New York Times and all the Disney characters that are still in copyright. It’s
the individual harm that bothers me: the struggling author whose words are
slurped up by the LLMs and regurgitated, potentially costing them readers; the
open source dev who chooses a copyleft licence on their work, only for an LLM
to dump complete copies of functions in someone else’s codebase with no
attribution or licence, stopping any future improvements from being released
for the public good.&lt;/p&gt;
&lt;p&gt;I can mostly mitigate those concerns in my personal use. When coding I generally
direct the agent closely enough that the opportunity to drop in massive chunks
of someone else’s code is negligible. I don’t use LLMs to generate big chunks
of written content. Overall I’m reasonably happy that I’m not making anyone
worse off.&lt;/p&gt;
&lt;p&gt;One related area that doesn’t affect me much because of my limited use, but does
concern me a lot, is corporate control. Most of the frontier LLMs are built by
big tech companies. As more and more people are using LLMs, they become an
avenue for control over information. It’s clear the people in charge should not
be trusted with that responsibility, as
&lt;a href=&#34;https://www.theguardian.com/technology/2025/may/14/elon-musk-grok-white-genocide&#34;&gt;Grok’s whole Boer War thing demonstrated&lt;/a&gt;.
I’m not sure what the answer is, though. Would a government built LLM be any
better? Probably not. The requirements to actually do the training are so high
that the only people who can do it are the ones you wouldn’t want to.&lt;/p&gt;
&lt;p&gt;So I’m left in the uncomfortable position of using tools by companies I don’t
trust, trained on data they shouldn’t have been trained on. Still, there can’t
be many more ethical issues, right?&lt;/p&gt;
&lt;h3 id=&#34;forced-features--flattery&#34;&gt;Forced Features &amp;amp; Flattery&lt;/h3&gt;
&lt;p&gt;Before we get onto the more obvious ethical topics, let’s go on a brief detour
so I can complain about how LLMs are being crowbarred into seemingly &lt;strong&gt;EVERY&lt;/strong&gt;
&lt;strong&gt;SINGLE&lt;/strong&gt; &lt;strong&gt;SERVICE&lt;/strong&gt;. I don’t want an LLM in my text editor. Or when shopping.
Or taking screenshots of everything my computer does every few seconds. I don’t
want LLMs anywhere outside a clearly-defined LLM box. With these forced
integrations, you’re even more in the dark about what their prompts are,
what information they have access to, what underlying model they’re using,
and so on.&lt;/p&gt;
&lt;p&gt;There’s no good reason for 90% of these implementations; there’s just some
weird Silicon Valley hype train that everyone is scared to miss out on. And
as a result we get more and more screen real estate taken up with things with
animated gradients that can’t be turned off and are prompted to talk in an
insufferable manner.&lt;/p&gt;
&lt;p&gt;The problem isn’t just that they’re everywhere, though, it’s how they behave
when they’re there. The overly sycophantic behaviour that most LLMs adopt can
hook users in the same way a social media app can fiddle with its algorithms to
encourage people to keep scrolling. It’s a lot harder to dismiss an LLM when
it’s blowing smoke up your ass. Now I don’t believe this behaviour is actually
intentionally nefarious, but more of an artifact of the models overfitting on
user approval instead of something useful like accuracy.&lt;/p&gt;
&lt;p&gt;The combination of going too far to please the user and hallucinating things
is incredibly problematic. How are ordinary, non-tech savvy, people meant
to deal with these models being forced in front of them? If you ask a question
and get back a flattering answer that’s roughly the shape of what you were
expecting, would you question it? Would a student, who was doing their homework
in Google Docs when an LLM inserted itself into their lives question it?&lt;/p&gt;
&lt;p&gt;This isn’t just a problem for naive users, either: I occasionally ask Claude
Code “Can we do X?” and discover much later that the answer should either be an
emphatic “NO”, or at least have a page of caveats attached. But the LLM wouldn’t
want to displease me, so it tries to do what’s asked of it, with all kinds of
weird hacks, and ultimately gets itself tied up in knots. This is not how a tool
should behave. If I wanted someone to say “Yes, and” repeatedly I’d join an improv troupe.&lt;/p&gt;
&lt;h3 id=&#34;environmental-expenses&#34;&gt;Environmental Expenses&lt;/h3&gt;
&lt;p&gt;The biggest ethical issue with LLMs is their environmental impact. Are we
adding more fuel to the dumpster fire that is our planet, just so my phone
can inaccurately summarise news headlines for me? The answer is a resounding
“maybe”. The companies making the LLM models are cagey about how much energy
goes into making them and then answering queries. There are estimates for
the latter, but there are several orders of magnitude between the most
optimistic and most pessimistic.
LLMs are definitely not &lt;em&gt;good&lt;/em&gt; for the environment. Basically no computer activity
is&lt;sup id=&#34;fnref:3&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:3&#34; role=&#34;doc-noteref&#34;&gt;3&lt;/a&gt;&lt;/sup&gt;. But you don’t see people campaigning to shutter YouTube (for example)
because of its environmental impact. I struggle to see the difference.&lt;/p&gt;
&lt;p&gt;I’m not saying considering the environmental impact isn’t important.
We are killing the planet. We are not doing enough about it. More attention
to the climate crisis in general is a good thing. But even with the most
pessimistic estimates on power usage, my personal LLM usage is a drop in
the ocean. I can have far more impact by reducing the amount I travel,
avoiding red meat, and so on.&lt;/p&gt;
&lt;p&gt;One issue I do have, though, is the forced integrations I mentioned. I’m happy
with &lt;em&gt;my&lt;/em&gt; personal usage. I’m not happy that a whole slew of LLMs are being
operated on my behalf. How much power is wasted by Google generating LLM
results in every search results page, versus how much benefit it provides?
How do people know if a search box will just use a few Watts doing a database
search, or burn kilowatts&lt;sup id=&#34;fnref:4&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:4&#34; role=&#34;doc-noteref&#34;&gt;4&lt;/a&gt;&lt;/sup&gt; spinning up an LLM and filling its context window?&lt;/p&gt;
&lt;p&gt;Then there’s the issue with training. Who knows what resources go into training
these models? Whether the amortised cost of that across all its users is worth
it? I think the answer to the first question is “only the tech companies making
the models”, and their utter silence on the second question says volumes.&lt;/p&gt;
&lt;p&gt;But this kind of industrial power isn’t unique to LLMs. There are many
industries that do far worse. And that’s not a justification, but I think it
makes sense to look at them more holistically. In an ideal world, where
we have functioning governments that listen to scientists&lt;sup id=&#34;fnref:5&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:5&#34; role=&#34;doc-noteref&#34;&gt;5&lt;/a&gt;&lt;/sup&gt;, we could address
all of this with legislation on emissions and clean power. Then all of these
industries will either manage to function in a manner compatible with keeping
the planet alive, or will have to stop. I appreciate that actually legislating
that is about as likely as waking up one day to find Claude has become an AGI,
though.&lt;/p&gt;
&lt;h3 id=&#34;slop--survival&#34;&gt;Slop &amp;amp; Survival&lt;/h3&gt;
&lt;p&gt;We’re nearly done, I promise. You can see why I didn’t want to try to tackle
these issues in an earlier post! The last big issue I want to talk about is
slop. It feels like it’s everywhere, and still getting worse.&lt;/p&gt;
&lt;p&gt;The Internet has slowly morphed from a wonderful place full of individual,
quirky sites, into a desert of bland corporate silos, and now into a wasteland
of AI slop. It’s becoming more and more difficult to find original information.
There’s this conspiracy theory called “Dead Internet Theory” that basically
says nearly all traffic on the Internet is generated by bots as part of a
co-ordinated effort to manipulate us. Obviously it’s unhinged, but at the
same time some elements of it are not that far off the mark. How long will it
be before the Internet is so drowned in slop that it is, effectively, dead?&lt;/p&gt;
&lt;p&gt;The other issue is what happens to the human jobs that have been replaced
with slop? It’s way too easy for a money-pinching publication to get rid of
their human staff in favour of pushing a few buttons and publishing the output.
You’d hope that market forces would balance things out, but there doesn’t seem
to be much sign of that happening. There are small movements of people putting
badges and other marks on their work to show it was made by a human, perhaps
that will take off?&lt;/p&gt;
&lt;p&gt;On a more personal level: what happens to software developers? As it stands,
code agents are nowhere near good enough to replace a human. At least not if
you want a non-trivial output that works, can be changed in the future, and
doesn’t have massive issues. Even if they get substantially better, I feel
like you’d still need a knowledgeable human in the loop to keep a rein on
everything. Does that mean the world needs far fewer software engineers?
Probably. But it doesn’t mean we’re all going to be out of jobs overnight
as some people seem to be predicting. The only way I can see that ever
happening is if there’s a radical shift in how software is made: a new language
of some kind that’s uniquely suited to LLMs, or new ways of composing software
together from smaller parts&lt;sup id=&#34;fnref:6&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:6&#34; role=&#34;doc-noteref&#34;&gt;6&lt;/a&gt;&lt;/sup&gt;.&lt;/p&gt;
&lt;p&gt;So… Yeah. Where does that leave us? I’ve written several thousand words about
how LLMs are an ethical mess, but I’m going to carry on using them. That’s uncomfortable.
But it’s not unlike using Amazon: they’re not a great company,
I’d rather give the money to a local shop or supplier, but I’m also not going
to massively inconvenience myself by being quixotic about it. Life is about
compromises like this, and it’d be silly to think otherwise. I’ll stand on
morals on some things&lt;sup id=&#34;fnref:7&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:7&#34; role=&#34;doc-noteref&#34;&gt;7&lt;/a&gt;&lt;/sup&gt;, but not to the extent of becoming a reclusive old
crank shouting at the clouds.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr/&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;That’s what &lt;a href=&#34;https://kagi.com/&#34;&gt;Kagi&lt;/a&gt; is for. It’s great. You should try it. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:1&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:2&#34;&gt;
&lt;p&gt;Or, maybe even worse, with added hallucinations. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:2&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:3&#34;&gt;
&lt;p&gt;Hell, basically no &lt;em&gt;human&lt;/em&gt; activity is good for the environment. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:3&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:4&#34;&gt;
&lt;p&gt;±several kilowatts, because who knows? &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:4&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:5&#34;&gt;
&lt;p&gt;and didn’t elect people who withdraw from climate agreements, re-open coal plants, and so on… &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:5&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:6&#34;&gt;
&lt;p&gt;I don’t &lt;em&gt;think&lt;/em&gt; I’m just saying this to reassure myself… &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:6&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:7&#34;&gt;
&lt;p&gt;I’ll grudgingly give Jeff Bezos money, but I sure as hell am not going to use any service from Elon Musk. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:7&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</content>
    </entry>
    <entry>
        <title>An app can be a ready meal</title>
        <link href="https://chameth.com/an-app-can-be-a-ready-meal/"/>
        <updated>2025-06-11T00:00:00Z</updated>
        <id>https://chameth.com/an-app-can-be-a-ready-meal/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/an-app-can-be-a-ready-meal/readymeal.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/an-app-can-be-a-ready-meal/readymeal.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/an-app-can-be-a-ready-meal/readymeal.jpg&#34; alt=&#34;A spaghetti carbonara ready meal, fresh out the microwave&#34; loading=&#34;lazy&#34; width=&#34;500&#34; height=&#34;335&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;It’s not a home-cooked meal, but it does the job sometimes.&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Three years ago I read “&lt;a href=&#34;https://www.robinsloan.com/notes/home-cooked-app/&#34;&gt;an app can be a home-cooked meal&lt;/a&gt;”
by Robin Sloan. It’s a great article about how Robin cooked up an app for his
family to replace a commercial one that died. It’s been stuck in my head ever since.
It’s only recently that I’ve actually done anything like Robin described,
though. Part of the reason was my brain got too hung up on the family aspect:
in my head, a home-cooked meal is one where your family or friends all gather
around to eat it with you (in much the same way as Robin’s app is used in
the article). It took me an embarrassingly long time to realise that you can
apply all the same arguments to an app built just for you. And it doesn’t even
have to be difficult. In fact, it can be more like a ready meal&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:1&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt; than a
family dinner.&lt;/p&gt;
&lt;h3 id=&#34;why-not-open-source&#34;&gt;Why not open source?&lt;/h3&gt;
&lt;p&gt;I love open source software. Almost everything I use day-to-day is open source,
and most things I write for myself I release as open source. I believe that
should be the default stance for most software. So why would you want to make
something and keep it just for yourself?&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;Even if an open source project garners no users whatsoever, there’s still some
pressure for it to meet certain standards&lt;sup id=&#34;fnref:2&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:2&#34; role=&#34;doc-noteref&#34;&gt;2&lt;/a&gt;&lt;/sup&gt;: if it needs configuring, then there
needs to be a mechanism for people to do that; the code needs to be of a
reasonable quality — my open source projects are linked from my CV, potential
clients might be looking at them — and there needs to be at least some attempt
at documentation or showing other people how to use it. There’s “building in
public” and then there’s “inviting the whole world to poke around your
drawer of shame”.&lt;/p&gt;
&lt;p&gt;The way I usually work on things that are mostly for me is I hack them together,
and then I gradually force myself to fix them up into a state where I consider
them acceptable for release. That stage isn’t fun, and sometimes it adds more
complexity than the whole original project. I’m not saying open source isn’t
worth it — far from it! — but there is definitely a balance to be struck
between the utility of making something open source and the amount of effort it
takes.&lt;/p&gt;
&lt;p&gt;For example, &lt;a href=&#34;https://chameth.com/home-automation-without-megacorps/&#34;&gt;I mentioned previously&lt;/a&gt;
that I wrote a home automation system. Hidden within that is a ~40 line
function that determines whether a fan should be turned on to keep a room cool.
The logic is entirely unique to the devices on my network and my requirements.
To try to put it into a form that would be useful to anyone but me would take
exponentially more effort. I’d basically be re-inventing Home Assistant, and
I explicitly started that project because I didn’t need that kind of complexity.
For a while I had this nagging feeling that I should find a way to open source
it. Then I had my&lt;sup id=&#34;fnref:3&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:3&#34; role=&#34;doc-noteref&#34;&gt;3&lt;/a&gt;&lt;/sup&gt; “lightbulb” moment: I could just &lt;em&gt;not&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Writing something entirely for yourself, without planning on open sourcing it
is surprisingly liberating: you don’t need to worry about documenting things;
you can change anything and everything at will without having to worry about
migration paths or whether it’ll break anyone’s workflow; you don’t have to
configure things, you can just code them how you want. Hell, you can hard-code
API keys right in the source code. Who cares!&lt;/p&gt;
&lt;h3 id=&#34;the-microwave-revolution&#34;&gt;The microwave revolution&lt;/h3&gt;
&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/an-app-can-be-a-ready-meal/is_it_worth_the_time.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/an-app-can-be-a-ready-meal/is_it_worth_the_time.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/an-app-can-be-a-ready-meal/is_it_worth_the_time.png&#34; alt=&#34;&amp;#39;Is It Worth the Time?&amp;#39; comic from XKCD. A table of &amp;#39;How often you do the task&amp;#39; vs &amp;#39;How much time you shave off&amp;#39;, with values showing how long you can work on making a task more efficient before spending more time than you save.&#34; loading=&#34;lazy&#34; width=&#34;571&#34; height=&#34;464&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;XKCD 1205&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;One of the other problems with making software just for yourself is that often
the time investment just isn’t worth it. If I spend a day automating something
that normally takes me five minutes, it’s going to take an awful long time to
“break even” on that time spent. If you’re open sourcing something then there
are ancillary benefits that may tip those scales: you might help other people,
have something to show off, etc. But if it’s just for you, then XKCD 1205 is a
harsh mistress.&lt;/p&gt;
&lt;p&gt;Recently, though, a new tool has emerged that tips the balance: LLM assistants.
&lt;a href=&#34;https://chameth.com/coming-around-on-llms/&#34;&gt;I’ve talked about my experience before&lt;/a&gt;. I’m definitely
not comfortable letting an LLM run roughshod over my published code, even if it
can help write some of it, but for private code? Why not! For a trivial example:
I have a USB key with an Arch Linux ISO on it, in case I need to troubleshoot
or reinstall my PC. Every so often I’ll update the ISO to the latest version.
It probably takes 5 minutes to do, so it’s probably only worth 25-50 minutes
to optimise it. Can I write, test, and debug a script to do it myself in that
time? Probably not. Can I get an LLM to generate some code to do it for me, and
then spend 10 minutes reviewing it&lt;sup id=&#34;fnref:4&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:4&#34; role=&#34;doc-noteref&#34;&gt;4&lt;/a&gt;&lt;/sup&gt;? Easily.&lt;/p&gt;
&lt;p&gt;To continue with the strained metaphor: LLMs are like the microwave that nukes
your ready meal. Pop in a prompt, let it use a load of power for a while, and
out pops your app. You can even do it on your phone while sat on a sofa&lt;sup id=&#34;fnref:5&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:5&#34; role=&#34;doc-noteref&#34;&gt;5&lt;/a&gt;&lt;/sup&gt;.
That changes the time equation even more: you can just squeeze in a prompt
whenever something comes to mind.&lt;/p&gt;
&lt;p&gt;This also goes back to the open source issue: I don’t feel comfortable
publishing something created by an LLM with minimal intervention on my part.
If I can prompt an LLM to do something, so can anyone else. We don’t need more
slop out there, and I don’t want there to be any confusion about what I’ve
written and what I’ve just prompted into existence. But for personal projects
that will never see the light of day it’s perfect.&lt;/p&gt;
&lt;h3 id=&#34;what-ive-made&#34;&gt;What I’ve made&lt;/h3&gt;
&lt;p&gt;Besides the home automation controller, I’ve built a constellation of smaller
tools: a script to arrange windows on my monitors how I want them&lt;sup id=&#34;fnref:6&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:6&#34; role=&#34;doc-noteref&#34;&gt;6&lt;/a&gt;&lt;/sup&gt;, one to
help me verify that my backups are working and can be restored, one to create
the files/folders for a new blog post, and some other odds and ends. They’re
all pretty small, and pretty specific to my setup.&lt;/p&gt;
&lt;p&gt;My biggest just-for-me project is a web app that I started to help me aggregate film recommendations.
It’s since morphed into a general personal data aggregation service: it deals
with data from GitHub, Todoist, Letterboxd, TMDB, Healthkit, and others. It also lets
me make re-orderable lists, store recipes, and more. Parts of this could
definitely be open sourced, and I might carve them out at some point, but it’s
mostly a glorious hodge-podge of things specific to me. Having all these
services in one place lets me make quick and dirty automations, for example:
when I create a Todoist note on my phone or watch, I often forget to set the
due date, so it doesn’t show up in the “Today” view. It was literally a few
lines of code to plumb things together so any inbox task without a due date
gets set to today automatically.&lt;/p&gt;
&lt;p&gt;If I was trying to write these things in a way that would be useful to other
people, or — for some of the features — without the aid of an LLM, I just
wouldn’t be bothered. It’d take too much time for questionable benefit. But
there’s a kind of joy in just being able to hack things together that work just
well enough for you. A ready meal will never compete with a home-cooked meal,
but sometimes it perfectly hits the spot.&lt;/p&gt;
&lt;hr/&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Image credits&lt;/th&gt;
&lt;th&gt;Creator&lt;/th&gt;
&lt;th&gt;Licence&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Spaghetti carbonara&lt;/td&gt;
&lt;td&gt;Wikimedia user Geni&lt;/td&gt;
&lt;td&gt;CC BY-SA 4.0&lt;/td&gt;
&lt;td&gt;&lt;a href=&#34;https://commons.wikimedia.org/wiki/File:Spaghetti_carbonara_ready_meal.JPG&#34;&gt;Wikimedia&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;XKCD 1205&lt;/td&gt;
&lt;td&gt;Randall Munroe&lt;/td&gt;
&lt;td&gt;CC BY-NC 2.5&lt;/td&gt;
&lt;td&gt;&lt;a href=&#34;https://xkcd.com/1205/&#34;&gt;XKCD&lt;/a&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr/&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;or a “TV dinner”, if you’re North-American-ly inclined. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:1&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:2&#34;&gt;
&lt;p&gt;Self-imposed pressure, for sure, but brains are going to do brain things. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:2&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:3&#34;&gt;
&lt;p&gt;very obvious in retrospect &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:3&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:4&#34;&gt;
&lt;p&gt;Reviewing LLM-generated code that shells out to &lt;code&gt;dd&lt;/code&gt; to write to the
root of a storage drive is perhaps the most intensely I’ve ever reviewed any
code to date. I really didn’t want to write a blog post about how an LLM
blatted my hard drive! &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:4&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:5&#34;&gt;
&lt;p&gt;Phones are terrible input devices for code, but they’re just about
passable for typing English. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:5&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id=&#34;fn:6&#34;&gt;
&lt;p&gt;I basically want a tiling window manager, but without all the effort and
weirdness of a tiling window manager. &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:6&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</content>
    </entry>
    <entry>
        <title>Coming around on LLMs</title>
        <link href="https://chameth.com/coming-around-on-llms/"/>
        <updated>2025-05-28T00:00:00Z</updated>
        <id>https://chameth.com/coming-around-on-llms/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/coming-around-on-llms/claude-hello.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/coming-around-on-llms/claude-hello.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/coming-around-on-llms/claude-hello.png&#34; alt=&#34;A screenshot of the Claude web UI with the prompt &amp;#39;Say &amp;#34;Hello!&amp;#34;&amp;#39;. The response is &amp;#34;Hello!&amp;#34;&#34; loading=&#34;lazy&#34; width=&#34;176&#34; height=&#34;172&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;Claude says hi.&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For a long time I’ve been a sceptic of LLMs and how they’re being used and
marketed. I tried ChatGPT when it first launched, and was totally underwhelmed.
Don’t get me wrong: I find the technology damn impressive, but I just couldn’t
see any use for it.&lt;/p&gt;
&lt;p&gt;Recently I’ve seen more and more comments along the lines of “people who
criticise LLMs haven’t used the latest models”, and a good number of developers
that I respect have said they use coding models in some capacity. So it seemed
like it was time to give them another shake.&lt;/p&gt;
&lt;p&gt;The first decision to make was which model to try. OpenAI are no longer the
only player in the game, every tech company of a certain size is now also
somehow an AI company&lt;sup id=&#34;fnref:1&#34;&gt;&lt;a class=&#34;footnote-ref&#34; href=&#34;#fn:1&#34; role=&#34;doc-noteref&#34;&gt;1&lt;/a&gt;&lt;/sup&gt;. I looked a bit at some benchmarks, and then mostly
ignored them and went with the only company that I didn’t outright hate:
Anthropic, and their model Claude.&lt;/p&gt;
&lt;h3 id=&#34;initial-impressions&#34;&gt;Initial impressions&lt;/h3&gt;
&lt;p&gt;The latest Claude models do feel a lot more “capable” than the earlier ChatGPT
versions I remember, but they also still have a lot of the same problems. At
their heart, they’re still text-prediction models, and still seem to be trained
to predict text that will please the user rather than be factually accurate or
useful.&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;One of the improvements that people kept mentioning was the ability for models
to access the web directly. I’ve seen it kick in a bit naturally, and asked
for it explicitly sometimes, and it’s… nothing special? Are people just really
bad at searching the web? Maybe they should try &lt;a href=&#34;https://kagi.com/&#34;&gt;Kagi&lt;/a&gt;? I
feel like about 80% of the time I could have found the information just as
quickly myself, about 10% of the time it searched and then hallucinated an
answer, and the remaining 10% it found something quicker than I otherwise
would have. Those aren’t great results, especially when you consider how
much cheaper and simpler just searching the web is.&lt;/p&gt;
&lt;p&gt;There is another way to trigger web searches in Claude: using “research mode”.
When you enable it, it splits off into different models: a “lead researcher” to
come up with a plan, and then some minions that execute it. It ends up doing
hundreds of web queries in the span of a few seconds. Suddenly I understand why
things like &lt;a href=&#34;https://anubis.techaro.lol/&#34;&gt;Anubis&lt;/a&gt; need to exist. I’ve not used
the “please launch a DoS attack” button since.&lt;/p&gt;
&lt;p&gt;What did impress me, though, was its ability to churn out reasonable-ish
code. It can hack together a bash script as well as I can, and do it far
faster than I’d be able to. Sometimes they even work. That made me wonder
what it would be like doing actual coding with it. Anthropic have a CLI tool
called &lt;code&gt;claude-code&lt;/code&gt;, so I paid them lots of money and gave it a spin.&lt;/p&gt;
&lt;h3 id=&#34;coding-with-claude&#34;&gt;Coding with Claude&lt;/h3&gt;
&lt;p&gt;The first thing I notice about &lt;code&gt;claude-code&lt;/code&gt; is that I really like the
interface. It’s basically an input box in a terminal. It doesn’t force me to
use a certain IDE or do things in a certain way. By default it asks before
making any changes, showing you a side-by-side diff of what it’s doing and
allowing you to provide feedback. If I had to design a way to interact with
a coding agent from scratch, I can’t think of many things I’d improve.&lt;/p&gt;
&lt;figure class=&#34;image full&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/coming-around-on-llms/code-session.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/coming-around-on-llms/code-session.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/coming-around-on-llms/code-session.png&#34; alt=&#34;A screenshot of claude-code. I ask it for the permalink for the latest post, and it gives a wrong answer. After prompting it again it gets it right.&#34; loading=&#34;lazy&#34; width=&#34;1121&#34; height=&#34;554&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;A simple example of a claude-code session&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The power in &lt;code&gt;claude-code&lt;/code&gt; versus just using the web UI is that it can use
tools. It can query &lt;code&gt;git&lt;/code&gt;, run &lt;code&gt;grep&lt;/code&gt; commands, even use &lt;code&gt;sed&lt;/code&gt; if it wants to
change something in lots of files at once. It’s very good at figuring out its
way around a codebase, even without any explicit instructions. You can see in
the screenshot that with a little prompting it managed to get the permalink
to this post; I didn’t tell it where the posts were stored, or how to work out
the latest, I just told it when it was wrong. If I’d run &lt;code&gt;/init&lt;/code&gt; before it
would have probably picked up on the fact that my posts have custom permalinks,
and noted it in &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;But how good is it at actually writing code? It’s like having a keen but not
particularly thorough Junior Engineer at your beck and call. If you give it
a clearly defined task and guidance on how to implement it (and maybe some
feedback as it suggests changes), it’s more than capable of doing it. If you
don’t give it enough guidance it tends to go more off the rails. I tried
having it generate a simple application from scratch with minimal technical
guidance and no review of what it was doing, and it made such a mess of it I
decided it was quicker to throw it away and start again by hand.&lt;/p&gt;
&lt;h3 id=&#34;the-man-behind-the-curtain&#34;&gt;The man behind the curtain&lt;/h3&gt;
&lt;p&gt;Even with sufficient guidance, at times it’s &lt;em&gt;really&lt;/em&gt; obvious that it’s an
LLM generating pleasing-token-strings and not something that genuinely
understands what it’s doing. It will spit out code like this:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-k&#34;&gt;if&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-nx&#34;&gt;err&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;!=&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-kc&#34;&gt;nil&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;{&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;chroma-k&#34;&gt;if&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-nx&#34;&gt;err&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-o&#34;&gt;==&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-nx&#34;&gt;sql&lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;.&lt;/span&gt;&lt;span class=&#34;chroma-nx&#34;&gt;ErrNoRows&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;{&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-w&#34;&gt;       &lt;/span&gt;&lt;span class=&#34;chroma-k&#34;&gt;return&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-kc&#34;&gt;nil&lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;,&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-nx&#34;&gt;err&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;}&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-w&#34;&gt;    &lt;/span&gt;&lt;span class=&#34;chroma-k&#34;&gt;return&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-kc&#34;&gt;nil&lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;,&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt; &lt;/span&gt;&lt;span class=&#34;chroma-nx&#34;&gt;err&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-p&#34;&gt;}&lt;/span&gt;&lt;span class=&#34;chroma-w&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Why’s that check for &lt;code&gt;sql.ErrNoRows&lt;/code&gt; there? It’s entirely pointless. There was
no instruction to check for it, none of the existing code checked for it, but
I assume it comes up quite a bit in the training data. But it didn’t
&lt;em&gt;understand&lt;/em&gt; why, so it put the check in, and returned the exact same thing as
if it hadn’t.&lt;/p&gt;
&lt;p&gt;It also occasionally tries to “cheat” or solve problems the wrong way. It
sometimes feels a bit like you’re asking for wishes from a Monkey Paw. “Stop
the unit tests failing”, you’ll say; Claude will respond with a request to
delete the failing test. The man behind the curtain isn’t particularly well
hidden, and the training data and pattern matching often shows through.&lt;/p&gt;
&lt;p&gt;This lack of understanding also makes it very hard to get Claude to use
comments in a sensible way. It absolutely loves doing nonsense like this:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-c1&#34;&gt;// render form
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-kd&#34;&gt;function&lt;/span&gt; &lt;span class=&#34;chroma-nx&#34;&gt;renderForm&lt;/span&gt;&lt;span class=&#34;chroma-p&#34;&gt;()&lt;/span&gt; &lt;span class=&#34;chroma-p&#34;&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;    &lt;span class=&#34;chroma-c1&#34;&gt;// ...
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&lt;span class=&#34;chroma-p&#34;&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Even with explicit instructions to avoid useless comments, or comments that only
explain what the code does (not why it does it), or other prompts. Again, I’m
chalking it up to a mixture of training data that does that, and lack of any
actual understanding about why a human developer may want a comment to exist.
Of course, there are plenty of flesh-and-blood devs out there that also don’t
comment effectively, so maybe I shouldn’t give Claude too much of a hard time
on this.&lt;/p&gt;
&lt;h3 id=&#34;-and-yet-&#34;&gt;… and yet …&lt;/h3&gt;
&lt;p&gt;So with all those problems, it sounds like it’s just not worth it, right?
Well, not quite. I feel like programming mostly consists of two distinct tasks:
thinking how to implement things; actually implementing them; and then debugging
why you’re off-by-one somewhere. Letting an LLM do the thinking is a big no-go
for me: it’s evidently not good at it, and frankly that’s one of the things
I most like about programming. Having let it loose on small projects, I shudder
to think what the codebases of all the “vibe coded” projects that are popping up
are like.&lt;/p&gt;
&lt;p&gt;For the second stage, though, it’s actually quite nice. If I’ve thought through
how I want something to be implemented, I can feed those steps to Claude and have
it churn out the otherwise not-too-interesting code. This requires far less
time and concentration on my part than writing code, to the extent that I can
be thinking about the next feature, or doing something else at the same time. I
don’t feel like I’m going to end up deskilling myself this way, as I’ve already
formed the idea of what I want to code, I’m just using the LLM to spit it out
faster than I can type it.&lt;/p&gt;
&lt;p&gt;A lot of the most boring bits of coding like implementing CRUD-y operations can
be summarised as “Look at this file/function. Do the same thing but slightly
differently elsewhere”. Claude is great at this, especially when explicitly
prompted like that. You’re basically playing to its pattern-matching strengths,
rather than asking it to come up with anything novel. Most of my prompts tend
to be prefixed with “Look at @some_file.go and @other_file.go.” to cue up the
patterns I want it to use, and I find this works well.&lt;/p&gt;
&lt;p&gt;As for debugging, it’s a mixed bag. It’s sometimes amazingly insightful, and
sometimes just runs around in circles trying the wrong things over and over.
It’s worth asking the question, but I definitely wouldn’t rely on it over my
own abilities.&lt;/p&gt;
&lt;h3 id=&#34;the-future&#34;&gt;The future&lt;/h3&gt;
&lt;p&gt;I’m probably going to carry on using &lt;code&gt;claude-code&lt;/code&gt;, at least for personal
projects. I’ve got so much done that I just wouldn’t have been
&lt;em&gt;bothered&lt;/em&gt; to do if I was doing it all by hand. I very much enjoy being in
the more “diffuse thinking” mindset, planning how things are going to work,
rather than being stuck in the mines digging out SQL queries. After all, who
wouldn’t want an over-eager assistant to work on all their hobby projects?&lt;/p&gt;
&lt;p&gt;Work is a slightly different matter: a private project or even an open source
project that disclaims any liability is different to something I’m being
paid to deliver, and bear responsibility for fixing if it’s not done correctly.
I’m not saying I won’t use it at all, but if I do it’ll be much more constrained
than I would in personal projects.&lt;/p&gt;
&lt;p&gt;As for non-code usages: I’m not sold. My sceptic hat is still firmly in place.
I don’t think chat is a particularly good interface for many things, and
hallucinations are still a big problem despite what people say. I hate the tide
of AI slop that’s taking over the Internet, and how LLM-powered chat
agents are being forced into every random product. You’re definitely not going
to be seeing any AI-generated blog posts from me!&lt;/p&gt;
&lt;p&gt;One topic I’ve not gone into here is the ethical concerns about using LLMs. They
obviously exist, and I do have thoughts, but that’s a topic for another day.&lt;/p&gt;
&lt;div class=&#34;footnotes&#34; role=&#34;doc-endnotes&#34;&gt;
&lt;hr/&gt;
&lt;ol&gt;
&lt;li id=&#34;fn:1&#34;&gt;
&lt;p&gt;I’m surprised it’s not gone more mainstream: why’s there
no Tesco Value LLM model, yet? &lt;a class=&#34;footnote-backref&#34; href=&#34;#fnref:1&#34; role=&#34;doc-backlink&#34;&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;
</content>
    </entry>
    <entry>
        <title>Debugging beyond the debugger</title>
        <link href="https://chameth.com/debugging-beyond-the-debugger/"/>
        <updated>2019-05-08T00:00:00Z</updated>
        <id>https://chameth.com/debugging-beyond-the-debugger/</id>
        <content xml:lang="en" type="html">&lt;figure class=&#34;image right&#34;&gt;
  &lt;picture&gt;
      &lt;source srcset=&#34;https://chameth.com/debugging-beyond-the-debugger/tools.avif&#34; type=&#34;image/avif&#34;/&gt;
      &lt;source srcset=&#34;https://chameth.com/debugging-beyond-the-debugger/tools.webp&#34; type=&#34;image/webp&#34;/&gt;
      &lt;img src=&#34;https://chameth.com/debugging-beyond-the-debugger/tools.jpg&#34; alt=&#34;Collection of tools hanging on a wall&#34; loading=&#34;lazy&#34; width=&#34;300&#34; height=&#34;396&#34;/&gt;
  &lt;/picture&gt;
  &lt;figcaption&gt;&lt;p&gt;Real-life debugging tools&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Most programming — and sysadmin — problems can be debugged in a
fairly straight forward manner using logs, print statements,
educated guesses, or an actual debugger. Sometimes, though, the
problem is more elusive. There’s a wider box of tricks that can
be employed in these cases but I’ve not managed to find a nice
overview of them, so here’s mine. I’m mainly focusing on Linux
and similar systems, but there tend to be alternatives available
for other Operating Systems or VMs if you seek them out.&lt;/p&gt;
&lt;h3 id=&#34;networking&#34;&gt;Networking&lt;/h3&gt;
&lt;h4 id=&#34;tcpdump&#34;&gt;tcpdump&lt;/h4&gt;
&lt;p&gt;&lt;code&gt;tcpdump&lt;/code&gt; prints out descriptions of packets on a network interface. You can
apply filters to limit which packets are displayed, chose to dump the entire
content of the packet, and so forth.&lt;/p&gt;
&lt;!--more--&gt;
&lt;p&gt;Typical usage might look something like:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# tcpdump -nSi eth0 port 80
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;tcpdump: verbose output suppressed, use -v or -vv for full protocol decode
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;listening on eth0, link-type EN10MB (Ethernet), capture size 262144 bytes
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;16:03:35.577781 IP6 2001:db8::1.54742 &amp;gt; 2001:db8::2.80: Flags [S], seq 2815779044, win 64800, options [mss 1440,sackOK,TS val 2378811665 ecr 0,nop,wscale 7], length 0
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;16:03:35.586853 IP6 2001:db8::2.80 &amp;gt; 2001:db8::1.54742: Flags [S.], seq 1522609102, ack 2815779045, win 28560, options [mss 1440,sackOK,TS val 3063610173 ecr 2378811665,nop,wscale 7], length 0
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;16:03:35.586877 IP6 2001:db8::1.54742 &amp;gt; 2001:db8::2.80: Flags [.], ack 1522609103, win 507, options [nop,nop,TS val 2378811674 ecr 3063610173], length 0
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;16:03:35.620678 IP6 2001:db8::1.54742 &amp;gt; 2001:db8::2.80: Flags [P.], seq 2815779045:2815779399, ack 1522609103, win 507, options [nop,nop,TS val 2378811708 ecr 3063610173], length 354: HTTP: GET / HTTP/1.1
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Here you can see the start of a plaintext HTTP request: the three-way
handshake as the TCP connection is established followed by a GET request.
Even if the data is encrypted as it will be in most cases, it’s often useful
to see the “shape” of the transmissions: did the client start sending data
when it connected, did the server ever respond, etc.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://danielmiessler.com/study/tcpdump/&#34;&gt;Daniel Miessler has a good tutorial on tcpdump&lt;/a&gt;
if you’re not familiar with it and don’t want to jump straight into the man
page.&lt;/p&gt;
&lt;h5 id=&#34;-with-docker&#34;&gt;… with Docker&lt;/h5&gt;
&lt;p&gt;Docker sets up separate network namespaces for each container. To see the
traffic across the interfaces of a single container you can &lt;code&gt;nsenter&lt;/code&gt; the
container’s network namespace:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# nsenter -t $(docker inspect --format &amp;#39;{{.State.Pid}}&amp;#39; my_container) -n tcpdump -nS port 80
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;This retrieves the PID for the container, and tells &lt;code&gt;nsenter&lt;/code&gt; to enter the
network (&lt;code&gt;-n&lt;/code&gt;) namespace from the given target (&lt;code&gt;-t&lt;/code&gt;) PID, and then run the
given command (in this case &lt;code&gt;tcpdump ...&lt;/code&gt;).&lt;/p&gt;
&lt;h4 id=&#34;openssl-s-client--s-server&#34;&gt;openssl s_client / s_server&lt;/h4&gt;
&lt;p&gt;When a connection is using TLS it’s often useful to try connecting to the
server and see what certificate it presents, algorithms it negotiates, and
so forth. OpenSSL offers two useful subcommands which can help with this:
&lt;code&gt;s_client&lt;/code&gt; for connecting as a client, and &lt;code&gt;s_server&lt;/code&gt; for listening to
connections.&lt;/p&gt;
&lt;p&gt;For example, using &lt;code&gt;s_client&lt;/code&gt; to connect to &lt;code&gt;google.com&lt;/code&gt; on the standard
HTTPS port shows us details about the server cert and its verification
status:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ openssl s_client -connect google.com:443
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;CONNECTED(00000003)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;depth=2 OU = GlobalSign Root CA - R2, O = GlobalSign, CN = GlobalSign
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;verify return:1
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;depth=1 C = US, O = Google Trust Services, CN = Google Internet Authority G3
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;verify return:1
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;depth=0 C = US, ST = California, L = Mountain View, O = Google LLC, CN = *.google.com
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;verify return:1
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;---
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Certificate chain
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt; 0 s:C = US, ST = California, L = Mountain View, O = Google LLC, CN = *.google.com
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;   i:C = US, O = Google Trust Services, CN = Google Internet Authority G3
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt; 1 s:C = US, O = Google Trust Services, CN = Google Internet Authority G3
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;   i:OU = GlobalSign Root CA - R2, O = GlobalSign, CN = GlobalSign
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;---
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Whereas connecting to my webserver and providing an unknown host in the SNI
field results in an SSL alert 112 (“The server name sent was not recognized”)
and no server certificate is sent:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ openssl s_client -connect chameth.com:443 -servername example.com
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;CONNECTED(00000003)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;140384831313024:error:14094458:SSL routines:ssl3_read_bytes:tlsv1 unrecognized name:../ssl/record/rec_layer_s3.c:1536:SSL alert number 112
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;---
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;no peer certificate available
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;---
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Often if you hit this kind of alert in an application the exact error will be
lost somewhere in the many layers between the SSL library and the logs, so
being able to directly connect and test can help diagnose a lot of issues.&lt;/p&gt;
&lt;p&gt;Once a connection is established you can read and write plain text and it
will be encrypted and decrypted automatically.&lt;/p&gt;
&lt;h4 id=&#34;java-apps&#34;&gt;Java apps&lt;/h4&gt;
&lt;p&gt;If a Java app is involved in the connection, you can enable a lot of built-in
debugging with a simple JVM property: &lt;code&gt;javax.net.debug&lt;/code&gt;. You can tweak
what exactly gets logged, but the easiest thing to do is just set the property
to &lt;code&gt;all&lt;/code&gt; and you’ll see information about certificate chains, verification,
and packet dumps:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ java -Djavax.net.debug=all -jar ....
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;found key for : duke
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;chain [0] = [
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;[
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  Version: V1
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  Subject: CN=Duke, OU=Java Software, O=&amp;#34;Sun Microsystems, Inc.&amp;#34;,
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  L=Cupertino, ST=CA, C=US
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;More information about Java’s debugging options is available on
&lt;a href=&#34;https://docs.oracle.com/javase/7/docs/technotes/guides/security/jsse/ReadDebug.html&#34;&gt;docs.oracle.com&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&#34;thread-and-core-dumps&#34;&gt;Thread and core dumps&lt;/h3&gt;
&lt;p&gt;Higher-level languages frequently provide an interactive way to dump the
current execution state of all of their threads (a “thread dump”). This
is useful to spot deadlocks, some types of race conditions, and as a
quick and dirty method of investigating hangs or excessive CPU usage.&lt;/p&gt;
&lt;p&gt;With both Java and Go applications you can send a QUIT signal to have a
thread dump printed out; Go applications will quit after doing so, Java
ones will carry on running. At most terminals you can hit &lt;code&gt;Ctrl&lt;/code&gt; and &lt;code&gt;\&lt;/code&gt; to
send a QUIT signal.&lt;/p&gt;
&lt;p&gt;For Java you can also use the &lt;code&gt;jstack&lt;/code&gt; tool from the JDK to dump threads
by PID; this can be useful if the application is running in the background
or has redirected sysout:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ jstack 8321
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Attaching to process ID 8321, please wait...
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Debugger attached successfully.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Client compiler detected.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Thread t@5: (state = BLOCKED)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt; - java.lang.Object.wait(long) @bci=-1107318896 (Interpreted frame)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt; - java.lang.Object.wait(long) @bci=0 (Interpreted frame)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt; - java.lang.ref.ReferenceQueue.remove(long) @bci=44, line=116 (Interpreted frame)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt; - java.lang.ref.ReferenceQueue.remove() @bci=2, line=132 (Interpreted frame)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt; - java.lang.ref.Finalizer$FinalizerThread.run() @bci=3, line=159 (Interpreted frame)
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;A core dump provides more complete information about the state of a process,
but is often more complex to interpret. The &lt;code&gt;gcore&lt;/code&gt; utility from GDB will
create a core dump of a process with a given PID. You can then generally
load the core file using your normal debugger, depending on the language
in question.&lt;/p&gt;
&lt;h3 id=&#34;system-calls&#34;&gt;System calls&lt;/h3&gt;
&lt;p&gt;&lt;code&gt;strace&lt;/code&gt; is the swiss army knife for seeing what a process is doing. It
details each system call made by a program (you can filter them down, of
course). For example:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ strace -e read curl https://google.com/
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;read(3, &amp;#34;\177ELF\2\1\1\0\0\0\0\0\0\0\0\0\3\0&amp;gt;\0\1\0\0\0 \236\0\0\0\0\0\0&amp;#34;..., 832) = 832
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;read(3, &amp;#34;\177ELF\2\1\1\0\0\0\0\0\0\0\0\0\3\0&amp;gt;\0\1\0\0\0P!\0\0\0\0\0\0&amp;#34;..., 832) = 832
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;read(3, &amp;#34;\177ELF\2\1\1\3\0\0\0\0\0\0\0\0\3\0&amp;gt;\0\1\0\0\0\200l\2\0\0\0\0\0&amp;#34;..., 832) = 832
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;read(3, &amp;#34;\177ELF\2\1\1\0\0\0\0\0\0\0\0\0\3\0&amp;gt;\0\1\0\0\0\20Q\0\0\0\0\0\0&amp;#34;..., 832) = 832
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;read(3, &amp;#34;\0\0\0\0\0\0\0\4\25\345\366\302\273sE6\365wI\225\321|\3435Z\362\216\372\215\251aO&amp;#34;..., 253) = 253
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&amp;lt;HTML&amp;gt;&amp;lt;HEAD&amp;gt;&amp;lt;meta http-equiv=&amp;#34;content-type&amp;#34; content=&amp;#34;text/html;charset=utf-8&amp;#34;&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&amp;lt;TITLE&amp;gt;301 Moved&amp;lt;/TITLE&amp;gt;&amp;lt;/HEAD&amp;gt;&amp;lt;BODY&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&amp;lt;H1&amp;gt;301 Moved&amp;lt;/H1&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;The document has moved
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&amp;lt;A HREF=&amp;#34;https://www.google.com/&amp;#34;&amp;gt;here&amp;lt;/A&amp;gt;.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;&amp;lt;/BODY&amp;gt;&amp;lt;/HTML&amp;gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;read(3, &amp;#34;\27\3\3\0!&amp;#34;, 5)                = 5
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;a href=&#34;http://www.brendangregg.com/blog/2014-05-12/strace-wow-much-syscall.html&#34;&gt;Brendan Gregg&lt;/a&gt;
has a nice guide on &lt;code&gt;strace&lt;/code&gt; and alternatives.&lt;/p&gt;
&lt;h4 id=&#34;-with-docker-1&#34;&gt;… with docker&lt;/h4&gt;
&lt;p&gt;When the application is running in docker you can usually just &lt;code&gt;strace&lt;/code&gt; it
from the host with the correct PID
(from e.g. &lt;code&gt;docker inspect --format &amp;#39;{{.State.Pid}}&amp;#39; my_container&lt;/code&gt;).
Sometimes you may need to trace the startup of an application though, which is
a bit trickier. Instead you can run a new container using the same PID
namespace as your target, and the permissions needed to &lt;code&gt;strace&lt;/code&gt;:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ docker run --rm -it --pid=container:my_container \
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  --net=container:my_container \
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  --cap-add sys_admin \
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  --cap-add sys_ptrace \
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;  alpine
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;From within the new container you can install strace, and trace any running
program within the target container using &lt;code&gt;strace -p&lt;/code&gt; as normal. To start a
new program you need access to the target container’s filesystem, which you
can get to via &lt;code&gt;/proc/1/root&lt;/code&gt; (PID &lt;code&gt;1&lt;/code&gt; being the main process that docker
started in the target container).&lt;/p&gt;
&lt;h3 id=&#34;files&#34;&gt;Files&lt;/h3&gt;
&lt;p&gt;Sometimes the problem might relate to file access. There are a couple of
straight forward — but nonetheless useful — tools which might help here.
&lt;code&gt;inotifywait&lt;/code&gt; uses the Linux &lt;code&gt;inotify&lt;/code&gt; subsystem to watch files or directories
for operations. For example:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ inotifywait -mr site/content
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Setting up watches.  Beware: since -r was given, this may take a while!
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;Watches established.
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;site/content/post/ MODIFY 2019-05-08-debugging-beyond-the-debugger.md
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;site/content/post/ OPEN 2019-05-08-debugging-beyond-the-debugger.md
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;site/content/post/ MODIFY 2019-05-08-debugging-beyond-the-debugger.md
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;site/content/post/ MODIFY 2019-05-08-debugging-beyond-the-debugger.md
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;site/content/post/ CLOSE_WRITE,CLOSE 2019-05-08-debugging-beyond-the-debugger.md
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Here the &lt;code&gt;-m&lt;/code&gt; switch makes &lt;code&gt;inotifywait&lt;/code&gt; monitor the files forever (instead
of exiting on the first modification, which is the normal behaviour) and &lt;code&gt;r&lt;/code&gt;
makes it recurse into the directory and monitor each file and subdirectory in
there.&lt;/p&gt;
&lt;p&gt;If you want to see what processes currently have a file open, &lt;code&gt;fuser&lt;/code&gt; is the
go-to tool. For example:&lt;/p&gt;
&lt;pre class=&#34;chroma-chroma&#34;&gt;&lt;code&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;$ fuser -v /
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;                     USER PID ACCESS COMMAND
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;/:                   root     kernel mount /
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;                     chris      2961 .rc.. systemd
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;                     chris      2986 .r... gdm-x-session
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;                     chris      2994 .r... dbus-daemon
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;                     chris      3001 .r... gnome-session-b
&lt;/span&gt;&lt;/span&gt;&lt;span class=&#34;chroma-line&#34;&gt;&lt;span class=&#34;chroma-cl&#34;&gt;# ...
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;h3 id=&#34;honourable-mentions&#34;&gt;Honourable mentions&lt;/h3&gt;
&lt;p&gt;These aren’t really debugging tools, but I feel it’s worth mentioning as
they often feature somewhere along the debugging-of-weird-problems journey.&lt;/p&gt;
&lt;p&gt;I’ve seen some weird and wonderful problems happen
because a disk is full, so a quick &lt;code&gt;df&lt;/code&gt; early on in the debugging process
never hurts. Some apps may hang, some may corrupt their config, some may
fall over and die; sometimes the manner in which they fail doesn’t obviously
point to a disk space issue.&lt;/p&gt;
&lt;p&gt;Another issue that comes up now and then — especially inside VMs or
other environment that don’t have a decent amount of “noise” happening —
is entropy exhaustion. A quick look at &lt;code&gt;/proc/sys/kernel/random/entropy_avail&lt;/code&gt;
should be enough to confirm that everything is ticking along nicely. If it’s
exceedingly low then you may find that anything involving random number
generation stalls (TLS connections for example).&lt;/p&gt;
</content>
    </entry>
</feed>
