<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:media="http://search.yahoo.com/mrss/"><channel><title><![CDATA[The Okami Project]]></title><description><![CDATA[The blog of a Sysadmin, Nerd and Internet Weirdo.]]></description><link>https://blog.okami.dev/</link><image><url>https://blog.okami.dev/favicon.png</url><title>The Okami Project</title><link>https://blog.okami.dev/</link></image><generator>Ghost 5.23</generator><lastBuildDate>Wed, 07 Oct 2026 22:50:32 GMT</lastBuildDate><atom:link href="https://blog.okami.dev/rss/" rel="self" type="application/rss+xml"/><ttl>60</ttl><item><title><![CDATA[Ceph, etcd, and the Sync-hole]]></title><description><![CDATA[Okami delves into a somewhat-known feature on Enterprise SSDs, and shows how much of an impact that they can have on your cluster performance.]]></description><link>https://blog.okami.dev/ceph-etcd-and-the-sync-hole/</link><guid isPermaLink="false">637c0e17b30cf40001742c99</guid><category><![CDATA[kubernetes]]></category><category><![CDATA[ceph]]></category><category><![CDATA[DevOps]]></category><category><![CDATA[testing]]></category><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Fri, 10 Dec 2021 01:19:27 GMT</pubDate><content:encoded><![CDATA[<p></p><p>Ceph and Etcd (stylized as <em>etcd) </em>are both fairly important pieces in the puzzle that is the cloud landscape. &#xA0;Ceph provides distributed, fault tolerant storage. Ditto for etcd, except that it&apos;s a database. Both use some pretty interesting ways of ensuring data integrity that can have surprising performance implications when you don&apos;t watch the hardware you&apos;re using.</p><p>This blog post aims to explain what these implications are, where they matter, and give some proof about these implications. Oh, and also provide some pretty mail-delivering foxes.</p><h2 id="how-storage-works-kinda">How storage works (kinda)</h2><p>Meet the storage fox. This storage fox is going to be our happy little metaphor to make explaining this a bit easier. Our storage fox delivers mail for us.</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2021/11/image-5.png" class="kg-image" alt loading="lazy" width="512" height="512"></figure><p>Lets say I write a letter (data) to my friend, Paul (the target storage device). I go and give this letter to the storage fox to deliver this mail. The storage fox gets a lot of these letters, so they chuck it into their bag (cache) to deliver <em>at some point</em> today. I can depend on the storage fox to deliver my letter; and I don&apos;t need to know when it gets delivered, just that it gets delivered at some point. This makes sending letters very quick, as I can just give it to the storage fox and continue with my day; knowing that the storage fox will get it to the destination.</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2021/11/image-6.png" class="kg-image" alt loading="lazy" width="512" height="512"></figure><p>There&apos;s a downside to keeping my mail in the bag. if Paul moves, or the fox gets hit by a truck, that mail isn&apos;t gonna get to the destination. It ends up lost.</p><p>This way of delivering mail is analogous to what asynchronous writes are within a storage system. I can tell the kernel to write some data, and it&apos;ll cache it (on the disk itself, maybe within RAM) while it&apos;s writing it to the final destination. My programs won&apos;t end up waiting for the storage operation to complete. This is great for situations where speed is key, and I don&apos;t <em>need</em> to guarantee it&apos;s been written. Asynchronous writes are also useful if I&apos;m going to be reading that data back, as the cache is often much faster to read from. That being said, the downside rears it&apos;s ugly head when the system loses power suddenly, or the storage disappears. </p><p>If you remember &quot;It is now safe to remove the USB&quot; in Windows, this is exactly why.</p><p>Now, lets say that I&apos;ve got a very important message to send to Paul. It&apos;s a cute cat photo I found on the internet, and I can&apos;t wait to see Paul&apos;s reaction. I give it to the storage fox, and I tell the fox to deliver this mail right away and tell me to come straight back when it&apos;s been delivered. This means I spend the entire time waiting for that mail to get to it&apos;s destination, but I will <em>know</em> that it&apos;s been delivered or not. Unfortunately that entire transaction takes much longer to complete, as I&apos;m now waiting for the fox to deliver the message, instead of carrying on with my day.</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2021/11/image-7.png" class="kg-image" alt loading="lazy" width="512" height="512"></figure><p>That&apos;s how a synchronous write works. I tell the kernel to ignore any caches and tell the storage target to <em>confirm</em> that the data has been fully written. Once I get that confirmation, I know that I can endure power loss, or unplug the device, and nothing will be lost. Synchronously writing data is slow though. If you&apos;ve noticed transfer speeds suddenly drop off a cliff when writing large files, you&apos;ve exhausted the cache and are now writing <em>directly</em> to the disk.</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2021/11/image-8.png" class="kg-image" alt loading="lazy" width="512" height="512"></figure><p><em>Why does that matter?</em></p><p>Writing data synchronously vs asynchronously is important. Writing asynchronously means you can achieve very fast IO per second (IOPS), but you haven&apos;t fully committed the data until another point later in time. Writing synchronously on the other hand is much slower; &#xA0;but you guarantee that a write is <em>actually </em>written. Most of the time, you don&apos;t really need to care, unless you&apos;re writing Databases or similar. This is where Ceph and etcd come in.</p><h2 id="ceph-and-etcd-like-safe-data">Ceph and Etcd like safe data.</h2><p>Moving on from the fox drawings, it&apos;s time to talk about Ceph and Etcd. They both write data synchronously... sorta. They use a concept known as write-ahead logging (WAL). This is essentially a log of what&apos;s going to get written somewhere. Data is first committed to the WAL (synchronously), and then to the persistent storage. </p><p>It&apos;s sorta like a cache, but fully explaining it (along with it&apos;s benefits) are outside the scope of this post. All you really need to be concerned with, is that the WAL is <em>always</em> written synchronously to ensure data integrity. This <em>kills </em>performance on a lot of drives; as they&apos;re most commonly built with asynchronous performance in mind.</p><p><em>Alright, I guess I get it now? Ceph and Etcd are slow, because they use synchronous writes, and synchronous writes are slow?</em></p><figure class="kg-card kg-image-card kg-width-wide kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2021/11/image-4.png" class="kg-image" alt loading="lazy" width="960" height="480"><figcaption>It gets worse.</figcaption></figure><h2 id="storage-can-lie-and-thats-sometimes-good">Storage can lie, and that&apos;s sometimes good.</h2><p>This is the heart of why I wrote this blog post. You can get your consumer storage devices; and enterprise drives. While consumer drives are measured in terms of raw speed; enterprise drives are measured in features and reliability. </p><p>Lets take a look at 2 drives, for example.</p><p>The first is a fairly standard Consumer drive, the Samsung 970 EVO Plus. This drive in particular is advertised as having sequential write speeds of 2300MB/s, and sequential read speeds of 3500MB/s. </p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2021/11/NVME1.jpg" class="kg-image" alt loading="lazy" width="3192" height="1564"></figure><p>The second drive is a Samsung SM953. With quoted speeds of 1,750 and 850 MB/s (read and write sequentially).</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2021/11/NVME2.jpg" class="kg-image" alt loading="lazy" width="3632" height="1176"></figure><p>By any aspect, the 970 EVO should be faster, right? It&apos;s advertised as being faster. </p><p>When writing synchronously, directly to the disk and bypassing the cache; the SM953 is way faster. This is because it lies.</p><p>In the above picture, the SM953 has little black boxes marked with the plus (+) symbol. These are capacitors that provide the SSD with power in the case of a power outage.</p><p><em>Doesn&apos;t that mean that even with asynchronous writes, the data is safe?</em></p><p>Yep, and the engineers who make these kinds of drives know that. When you force a synchronous write to the drives with these capacitors on, these drives ignore that and treat it as if it were an asynchronous request. They lie that they&apos;re actually writing the data to the storage, since they can tolerate power loss. It&apos;s not 1:1 the same, but it&apos;s close enough to be negligible. </p><p>That makes these kinds of drives faster for applications like Ceph and Etcd; than traditionally &apos;faster&apos; drives. This is because the WAL (cache) gets written synchronously, which these drives are <em>really</em> good at.</p><p>It&apos;s key to note that I will be referring to drives with the power loss protection feature as PLP drives.</p><p>One other thing that&apos;s worth nothing, is that enterprise drives are rated for a <em>far</em> higher level of endurance than typical SSDs. They often outlast their use case, so you can frequently see them on ebay for cheap; with only ~5% of their total life used.</p><h2 id="the-testing-set-up">The testing set-up</h2><p>Testing this wasn&apos;t &#xA0;easy. I figured I could come out with metrics and say &quot;x drive is 30x faster than another&quot;; but that&apos;s really, really hard to actually prove. Different workloads are better with different drives.</p><p>To make things simple(r), I set out to to try and prove the following hypotheses:</p><blockquote>Direct, synchronous writes are faster on drives with power loss protection</blockquote><blockquote>Ceph, and (to a lesser extent) etcd perform better on drives with power loss protection</blockquote><p>Where &apos;faster&apos; and &apos;better&apos; can be considered as that higher IOPS, higher bandwidth and lower latency in a like-for-like comparison. This leaves some interpretation up to the reader, &#xA0;but I&apos;m going to consider <em>any</em> comparison being beyond a 10% margin to be &apos;faster&apos;.</p><p>As I&apos;m in the process of a cluster rebuild, I put one of my hypervisors to work being the testing platform.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2021/11/PXL_20211101_182533712.jpg" class="kg-image" alt loading="lazy" width="4032" height="3024"><figcaption>Pictured, the test dummy. Complete with 4 NVME SSDs.</figcaption></figure><p>To ensure that the tests are a like-for-like comparison, they are being run on the same host, with the same OS, the same packages and mostly the same variables.</p><p>One of the considerations I had when thinking about how to test this, was whether the host IO (i.e the install of the OS) could negatively affect performance. To solve this, I&apos;m running Alpine Linux entirely in memory. This eliminates any considerations of &apos;host OS storage&apos;; as it&apos;s just RAM.</p><p>I made a number of other considerations like the above, but the basic gist is that I tried my best to keep this as small, fast and fair as possible.</p><p>The plan was to run a set of 3 tests. </p><ul><li>Raw (FIO based) benchmark published by ETCD</li><li>The etcd benchmarking tool, targetting a formatted filesystem</li><li>A disk benchmarking tool intended for use in Kubernetes.</li></ul><p>For each of these tests, I put 4 drives up against the task.</p><p>The four drives are as follows:</p><ul><li><code>nvme0n1</code> Samsung PM963 (960GB)</li><li><code>nvme1n1</code> Samsung 970 EVO Plus (500GB)</li><li><code>nvme2n1</code> Samsung SM953 (480GB)</li><li><code>nvme3n1</code> Samsung 970 EVO Plus (250GB)</li></ul><p>The first (PM963), and third (SM953) drives have power loss protection. </p><p>These are the ones I was hoping to prove are faster.</p><p>Lets begin!</p><h2 id="test-1-fio">Test 1: FIO</h2><p>After getting the OS booted, I ran FIO on each drive on a mounted directory under <code>/mnt/</code>. </p><p>This first test was looking at proving that the latency is lower on PLP drives. </p><p>This test was almost like-for-like copied from a post about etcd performance.<sup>[1]</sup></p><p>The FIO command run was as follows:</p><pre><code>export DISKNAME=nvme0n1
fio --rw=write --ioengine=sync --fdatasync=1 --directory=/mnt/${DISKNAME}p1/testdir --size=22m --loops=10 --bs=2300 --name=$DISKNAME</code></pre><p>(The referenced blog post goes into great depths talking about the WAL and etcd. I highly recommend checking it out.)</p><p>The results were promising.</p><!--kg-card-begin: html--><iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/fio.html" loading="lazy"></iframe><!--kg-card-end: html--><p>Please note, for this graph, the results are not <em>linear. </em>They&apos;re logaritmic, as the difference between both sets of drives are massive. (~30 usec vs 785, for example)</p><p>Comparing the average sync latencies, you can see that the latency for each percentile of the test was <em>lower</em> on drives with power loss protection. </p><p>In the graphs (and all subsequent graphs), the PLP drives are the ones that begin with MZ. </p><p>Since the drives are able to lie about the disk writes, they&apos;re able to respond <em>much</em> quicker than the drives that remain truthful.</p><h2 id="test-2-etcd">Test 2: Etcd</h2><p>Moving on, I installed etcd onto the host, and launched it, using each drive as the target &apos;data&apos; directory.</p><p>The actual command run was as follows:</p><pre><code>export DISKNAME=nvme0n1 
etcd --force-new-cluster --data-dir $ETCDDIR/data --log-outputs /var/log/${DISKNAME}
</code></pre><p>I then used the <a href="https://github.com/etcd-io/etcd/tree/main/tools/benchmark">etcd benchmarking tool</a> to run my tests on the freshly created etcd cluster.</p><pre><code>benchmark \
	--target-leader \
    --conns=100 \
    --clients=1000 \
    put \
    --key-size=8 \
    --sequential-keys \
    --total=100000 \
    --val-size=256</code></pre><p>Since I wanted to really stress these drives, I opted for a higher number of connections, clients and keys; than is what&apos;s usually recommended for benchmarking.</p><p>The results are as follows:</p><!--kg-card-begin: html-->
<iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/etcd.html" loading="lazy"></iframe><!--kg-card-end: html--><p>From the 10% to the 75% percentile, the results look almost exactly the same. Essentially, there&apos;s almost no difference in 75% of writes. Beyond that is where you start to see a difference in results, with the PLP drives increasing in latency at a much more gradual pace than the non-PLP drives. The difference is nowhere near the same as the fio based tests, and I suspect that etcd may not be immediately syncing writes to disk, or that a bottleneck/slowdown might be elsewhere. I didn&apos;t want to dwell too far on this, as there&apos;s still a fairly clear difference in the higher percentiles, and I think that demonstrates &apos;better&apos;.</p><h2 id="test-3-ceph">Test 3: Ceph</h2><p>Ceph was setup by running a single node Kuberetes cluster and deploying <a href="https://rook.io/">Rook</a> onto it. From there, I was able to create one cluster for each drive fairly painlessly.</p><p>I built the cluster using the following Kubernetes components.</p><pre><code>kubernetes deployment method: kubeadm
Container runtime interface: cri-o
Container runtime: crun
Network layer: kube-router (replacing kube-proxy)</code></pre><p>Once Ceph and rook was installed, I tinkered around with modifying the number of replicas and OSDs for each cluster; and ran tests for each change. </p><p>If you&apos;re totally new to Ceph, the replica count is the number of times one &apos;thing&apos; will get written to a disk. A replica 3 setup will mean that 3 replicas of one &apos;thing&apos; are stored when you issue a write.</p><p>An OSD, is a &apos;partition&apos; of a disk. For faster drives, more OSDs can increase your performance. If you&apos;re new to Ceph, just think of it as different ways to configure your cluster. Some are faster (lower replicas), and some change how your data ends up on the disk (OSDs). </p><p>I ended up running 4 rounds of tests, which are as follows.</p><ul><li>4 OSDs, 1 Replica</li><li>4 OSDs, 3 Replicas</li><li>5 OSD, 1 Replica</li><li>5 OSD, 3 Replicas</li></ul><p>This would give me a fairly comprehensive set of data to work with, and also allow me to see how best to deploy things when the time comes to bring up my new infra.</p><p>Once I figured out the different cluster types I&apos;d test, I decided to use the <code>yasker/kbench</code> image to benchmark, as it seems to give a pretty nice set of results. It also seems fairly &apos;purpose built&apos; for Kubernetes, so this will hopefully give results that closely align with what I&apos;d expect to see in a &apos;real world&apos; deployment.</p><figure class="kg-card kg-code-card"><pre><code class="language-yaml">---
apiVersion: batch/v1
kind: Job
metadata:
  name: fio
spec:
  backoffLimit: 4
  template:
    spec:
      restartPolicy: Never
      containers:
      - name: dbench
        image: yasker/kbench:latest
        imagePullPolicy: Always
        # privilege needed to invalid the fs cache
        securityContext:
          privileged: true
        env:
          - name: SIZE
            value: 20G
          - name: FILE_NAME
            value: /data/testfile
        volumeMounts:
          - name: disk
            mountPath: /data
      volumes:
        - name: disk
          persistentVolumeClaim:
            claimName: fio
            </code></pre><figcaption>The Job YAML for each test.</figcaption></figure><p>With that out of the way, lets look at the results.</p><p><strong>Bandwidth</strong></p><p>First up, is bandwidth. This is the <em>amount</em> of data that is able to get onto the disks, measured in Kilobytes per second.</p><!--kg-card-begin: html--><iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/random_write_Bandwidth.html" loading="lazy"></iframe><!--kg-card-end: html--><p>These tests are for random writes. This is the sort of workload you&apos;d encouter if you were updating a game, or your OS. Lots of files, in a number of different places. This is often where hard drives <em>suck</em> because they need to move the physical R/W head to find the data.</p><!--kg-card-begin: html--><iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/sequential_write_Bandwidth.html" loading="lazy"></iframe><!--kg-card-end: html--><p>This graph displays the sequential writes over 4 different cluster configurations. A sequential write is essentially &apos;one after the other&apos;. A real example of this would be streaming a movie. </p><p>In both cases, the power loss protection drives were <em>far</em> ahead of the competition. We can also see that the number of OSDs matter less in sequential writes, and this is due to the fact that 4 streams to the same drive vs 5 streams isn&apos;t going to matter when you&apos;re writing large files. What does seem to matter, is the number of replicas. </p><p><strong>IOPS</strong></p><p>IOPS is an important metric. It essentially dictates how many operations per second a drive can complete. This correlates with latency, as a higher latency will result in lower IOPS (and vice versa).</p><!--kg-card-begin: html-->

<iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/random_write_IOPS.html" loading="lazy"></iframe><!--kg-card-end: html--><p>The random results above paint a similar picture to bandwidth. PLP drives are scoring far ahead of their counterparts. The same use cases apply for random and sequential; so a high random IOPs metric will mean your applications spend less time waiting.</p><!--kg-card-begin: html--><iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/sequential_write_IOPS.html" loading="lazy"></iframe><!--kg-card-end: html--><p>Next, sequential IOPS. These also show that when writing sequentially, you&apos;re not going to encounter as much of a bottleneck when using between 4 and 5 OSDS.</p><p>So far, we have a 2/2 winning set landing solely at the PLP drives. Next up is latency.</p><p><strong>Latency</strong></p><p>Latency is a figure you want to keep low for a storage system. Much like how waiting in line for the poor employee to scan your meal deal feels bad; data waiting to be written feels bad. &#xA0;</p><!--kg-card-begin: html--><iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/random_write_Latency.html" loading="lazy"></iframe><!--kg-card-end: html--><p>Here we can see the latency (in nanoseconds) for the random writes is nearly 3x slower when writing to non PLP drives.</p><!--kg-card-begin: html--><iframe width="100%" height="500em" src="https://okamidash.github.io/sync-results/random_write_Latency.html" loading="lazy"></iframe><!--kg-card-end: html--><p>The same can be said for sequential writes. </p><p>That&apos;s 3 for 3. Looks like it&apos;s pretty clear that PLP drives are <em>far</em> better suited for Ceph, than non PLP drives.</p><h2 id="limitations">Limitations</h2><p>Before I get into the conclusion, it is worth raising a limiting factor in my testing. </p><blockquote>I do not have a large sample set</blockquote><p>This is true. I have only been testing on four drives. It&apos;s something I can&apos;t exactly get around without a research budget, and my spare salary is currently spent on other adventures. To combat this, I am going to be publishing all of my testing assets on <a href="https://github.com/okamidash/sync-results">Github</a>, and if you have any spare time, please run the tests yourself and submit a pull request. </p><figure class="kg-card kg-bookmark-card"><a class="kg-bookmark-container" href="https://github.com/okamidash/sync-results"><div class="kg-bookmark-content"><div class="kg-bookmark-title">GitHub - okamidash/sync-results</div><div class="kg-bookmark-description">Contribute to okamidash/sync-results development by creating an account on GitHub.</div><div class="kg-bookmark-metadata"><img class="kg-bookmark-icon" src="https://github.com/fluidicon.png" alt><span class="kg-bookmark-author">GitHub</span><span class="kg-bookmark-publisher">okamidash</span></div></div><div class="kg-bookmark-thumbnail"><img src="https://opengraph.githubassets.com/3a696cd3be095b025c28ebe8400c91866c11410c8bf046428018544125e23d2d/okamidash/sync-results" alt></div></a></figure><h2 id="conclusion">Conclusion</h2><p>Storage is hard. There is a lot of information to take in, and I think it&apos;s important to note that I do not know all the answers. This isn&apos;t my specialty. I&apos;m willing to accept that I am wrong on a number of these things.</p><p>Some key points to note having said that. These tests are a like-for-like comparison between drives. I attempted to control every condition that may have affected the performance otherwise. By all accounts, the testing I have performed is valid; but don&apos;t take it as gospel.</p><p>I wrote this blog post to prove 2 key points. </p><blockquote>Direct, synchronous writes are faster on drives with power loss protection</blockquote><p>I feel like I have been able to prove this fairly well. Looking at the results from section 1 (FIO), it&apos;s fairly clear that the drives perform better, scoring a lower latency; when given a direct, synchronous workload.</p><blockquote>Ceph, and (to a lesser extent) etcd perform better on drives with power loss protection</blockquote><p>The results from Ceph I feel are the shining star in all of this. In every test, with varying degrees of replication, OSDs, and workloads (Sequential and Random), the drives equipped with power loss protection scored <em>significantly</em> higher than the ones without. That&apos;s about as clear cut as it gets. </p><p>If you want to go ahead with deploying Ceph on non-plp drives, I would suggest a replica 1 setup, with either 4 or 5 OSDs for each drive. You ideally want to minimize the number of total writes to the drive for each &apos;user&apos; write. Backwards logic, but the tests seem to show that you&apos;ll acheive higher performance when you can write <em>less</em>. </p><p>Anyway. </p><p>This is been one of the most difficult posts I&apos;ve written in a long time. I&apos;m currently staring at the clock, it&apos;s 1:15 AM on a Friday. Countless weekend have gone into this to give the level of detail I think this work deserves. I would <em>love</em> to get your feedback.</p><p>Hopefully this will guide your drive purchasing in the future if you&apos;re deploying Ceph and/or etcd. </p><p>Signing off, </p><p>Okami.</p><p>(Btw, the art is made by my soon-to-be wife. Thanks Aur &lt;3 )</p><p></p><h2></h2>]]></content:encoded></item><item><title><![CDATA[Production K8s, one year on.]]></title><description><![CDATA[<p>It&apos;s been one year since I deployed hydra, my production Kubernetes cluster.</p><p>This post is going to be a write-up of the things I did right, and the things I did wrong, warts and all. I&apos;ll close off by talking about my plans for the future.</p>]]></description><link>https://blog.okami.dev/one-year-of-hydra/</link><guid isPermaLink="false">637c0e17b30cf40001742c98</guid><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Sat, 12 Jun 2021 17:02:34 GMT</pubDate><media:content url="https://blog.okami.dev/content/images/2021/06/Screenshot-2021-06-12-at-11-16-09-hydra---Kubernetes-Dashboard.png" medium="image"/><content:encoded><![CDATA[<img src="https://blog.okami.dev/content/images/2021/06/Screenshot-2021-06-12-at-11-16-09-hydra---Kubernetes-Dashboard.png" alt="Production K8s, one year on."><p>It&apos;s been one year since I deployed hydra, my production Kubernetes cluster.</p><p>This post is going to be a write-up of the things I did right, and the things I did wrong, warts and all. I&apos;ll close off by talking about my plans for the future.</p><figure class="kg-card kg-image-card kg-width-full kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2021/06/image-4.png" class="kg-image" alt="Production K8s, one year on." loading="lazy" width="1407" height="150"><figcaption>I had a calendar entry for this :p</figcaption></figure><h2 id="an-unhealthy-pod-is-a-dead-pod">An unhealthy pod is a dead pod</h2><p>In the old world, if I deployed a virtual machine, or a container; they would become my <em>pet</em>. It was my responsibility to manage, it was my responsibility to treat with care and affection. I gave it a name and I wouldn&apos;t dare delete it. In Kubernetes, that&apos;s not how things work. The &apos;closest&apos; [1] analogue to a container is a Pod. How do you restart the pod if the process fails to start properly? You <em>don&apos;t</em>. </p><p>Pods get removed. They get <em>deleted</em>. They&apos;re not pets anymore, not here.</p><p>There isn&apos;t time to name your pets. They&apos;re cattle now. If one cow dies, whatever. If the whole herd does? you&apos;ve got a problem. This kind of thinking is what I had to switch to when I moved to Kubernetes. &#xA0;You don&apos;t restart pods. You destroy them.</p><p>I think the greatest realization of this was when <a href="https://blog.okami.dev/losing-4-days-in-4-seconds/">I lost my <em>lovingly</em> crafted dashboards</a> because I deleted a pod. </p><p>This has greater implications than just deleting pods. Any changes you want to make permanent? It needs to be mutable either with a PVC or Configmap or whatever. This forces you to think more carefully about what files and settings are <em>important, </em>because if you just mount the entire container as a PVC, you won&apos;t be able to reap the benefits of updating. If you don&apos;t mount the right files, when the pod gets deleted, you&apos;ve lost those changes forever. This kind of understanding is made a lot easier with documentation from the developers of the service you&apos;re deploying; but sometimes you find that some edge cases (my Nextcloud deployment has some CM mounts in the /etc dir to increase the PHP filesize, for example).</p><p>[1] One could probably argue that a container (within a pod) is the closest to a container within docker, but a pod is the lowest &apos;level&apos; one works within Kubernetes.</p><h2 id="get-your-secrets-right-from-the-start">Get your secrets right from the start</h2><p>Secrets, API keys, passwords, all of that. It&apos;s sensitive (obviously). Make sure the management of secrets is done in a sane way from day 0. &#xA0;I made the mistake of deploying stuff using secrets stored in git. While secrets aren&apos;t encrypted in Kubernetes[2], it&apos;s an incredibly easy method of attack if your secrets are stored in CI.</p><figure class="kg-card kg-image-card kg-width-full kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2021/06/image-8.png" class="kg-image" alt="Production K8s, one year on." loading="lazy" width="1332" height="885"><figcaption>I am not a smart person.</figcaption></figure><p>This is the biggest criticism I can make of my own infrastrucutre. I didn&apos;t do secrets properly, and as a result they&apos;re scattered about within my hydra repo. Anyone that gets granted access to the repo, gets the keys to the kingdom.</p><p>One of the tasks on my TODO list is to remove the secrets entirely; but it&apos;s something I should have avoided from the start.</p><p>[2] <a href="https://twitter.com/okami_dash/status/1396605990104145927">Base64 is not encryption.</a></p><h2 id="status-checking-is-only-as-good-as-the-check">Status checking is only as good as the check</h2><p>Status checking provides a great insight into the health of a cluster, from a user perspective. It&apos;s a nice, blackbox way of saying &quot;Is everything running fine?&quot;. </p><p>The way my network and DNS is designed means that internally my services can function absolutely fine; but externally it&apos;s fucked. This makes status checking a little bit difficult, since I can&apos;t rely on checking from within Kubernetes. &#xA0;</p><p>To solve this, I use uptime robot; it&apos;s not ideal since there&apos;s a pretty hefty delay between knowing the status of something, and uptime robot updating.</p><p>That being said, having these checks is still useful. Knowing the status of your estate is one of the most important things to have when you&apos;re running distributed services.</p><h2 id="use-a-dedicated-os-for-the-cluster">Use a dedicated OS for the cluster</h2><p>Currently my nodes use a mixture of <a href="https://kubic.opensuse.org/">OpenSuse Kubic</a> and <a href="https://getfedora.org/">Fedora</a>. I started with Fedora, and over time i&apos;ve transitioned to Kubic; due to the ease of use.</p><blockquote>Ease of use? What do you mean?</blockquote><p>Think about maintaining a virtual machine. Keeping it updated. Making sure that config doesn&apos;t change. Making modifications. Now multiply that by 9. It&apos;s not fun. Even using tools like Ansible to automate most of it. Especially on version based operating systems that release fairly frequently.</p><figure class="kg-card kg-image-card kg-width-full"><img src="https://blog.okami.dev/content/images/2021/06/image-7.png" class="kg-image" alt="Production K8s, one year on." loading="lazy" width="1045" height="589"></figure><p>Using a dedicated operating system for the job that&apos;s purpose built for Kubernetes has saved a number of headaches. </p><p>I don&apos;t need to worry about cgroup configuration, repo setup, all of that. It&apos;s not a <a href="https://github.com/oxide-one/haikoo/tree/main/roles/kubernetes/prepare_node/tasks">particularly easy task</a> to prepare a node for use in a cluster, and maintaining the OS on top of the cluster itself isn&apos;t a particularly fun thing either. Keeping Fedora nodes in the cluster makes sense if your nodes are <em>pets</em>, since you can afford to maintain them. My nodes aren&apos;t pets. They&apos;re cattle. They&apos;re part of the system.</p><h2 id="avoid-dependency-locks">Avoid dependency locks</h2><p>When building distributed systems, you will have dependencies. Don&apos;t make those dependencies circular. I intentionally built my cluster to avoid these circular locks.</p><p>Example (albeit not a great one):</p><blockquote>You set your controlplane endpoint to be &apos;api.example.com&apos;, and put the IP in /etc/hosts temporarily.<br>You deploy the cluster, and then deploy the DNS that provides the NS authority for &apos;example.com&apos;.<br>You&apos;re then relying on the DNS being hosted by Kubernetes, to provide Kubernetes critical function.</blockquote><p>Don&apos;t do this. It will result in <em>something</em> being cluster-external; but that&apos;s a tradeoff that is absolutely worth it. Soft dependencies are easier to deal with, such as deploying a CI system to deploy changes to the cluster itself; as you can still do things manually if it fails.</p><h2 id="it-never-stops">It never stops</h2><p>As is the case in a DevOps world, the work never stops. It will never be done. The mental laundry list of things I want to change and fix grows bigger with every day. That&apos;s the reality of Kubernetes, and it&apos;s one I&apos;ve learnt over the past year of maintaining a kubernetes cluster in production.</p><h2 id="the-future">The future</h2><p>My current stack sits ontop of oVirt; which is fantastic. I&apos;ve noticed that the development is a bit slow, and a lot of resources has gone into other RH products; so i&apos;m considering the replacements.</p><p>The one I&apos;m planning (and currently developing); is a dual stack cluster. This is a Kubernetes in Kubernetes cluster using KubeVirt. Taking over the key points highlighted above, the thing I&apos;m most eager to work with, is secrets. Once the design is more concrete, I&apos;ll write it up. Don&apos;t think anyone has gotten KinKy before :p</p>]]></content:encoded></item><item><title><![CDATA[Using Helm with Kustomize effectively.]]></title><description><![CDATA[A look into the use of the HelmChartInflationGenerator to make Helm play nice with Kustomize.]]></description><link>https://blog.okami.dev/kustomizing-with-helm/</link><guid isPermaLink="false">637c0e17b30cf40001742c97</guid><category><![CDATA[kubernetes]]></category><category><![CDATA[DevOps]]></category><category><![CDATA[Kustomize]]></category><category><![CDATA[Helm]]></category><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Sun, 28 Mar 2021 00:19:14 GMT</pubDate><content:encoded><![CDATA[<p><a href="https://kustomize.io/">Kustomize</a> and <a href="https://helm.sh/">Helm</a> are both effective tools for managing <em>things</em> within Kubernetes. I am a big fan of Kustomize, and use it to manage my Kubernetes cluster, <em>hydra.</em></p><p>Although I primarily use Kustomize, I maintain a number of services that have a recommended deployment option of Helm, so what I would do is template out files to disk and then use Kustomize to merge and deploy them. This gives the benefit of being able to <code>kustomize build</code> <em>everything</em> within my cluster, instead of doing <code>helm install</code> for a number of deployments. </p><p>Since I try to stick to kustomize, I would often template out Helm deployments to a file, then deploy with kustomize. This was the going method for some of my deployments, until the <code>helmChartInflationGenerator</code> was added to kustomize.</p><p>The <code>helmChartInflationGenerator</code> allows me to use a Helm chart as a source, and directly use it from Kustomize. It has a number of options, but it&apos;s pretty simple to use. Lets go through an example with deploying <a href="https://haproxy-ingress.github.io">haproxy-ingress</a>.</p><h2 id="example-deploying-haproxy-ingress">Example: Deploying Haproxy Ingress</h2><p>Lets start by creating a <code>kustomization.yaml</code></p><pre><code class="language-YAML"># kustomization.yaml

# Tell kustomize to use the generator in chartInflator.yaml
generators:
- chartInflator.yaml


</code></pre><p>Next, we&apos;ll write the <code>chartInflator.yaml</code>.</p><pre><code># chartInflator.yaml

apiVersion: builtin
kind: HelmChartInflationGenerator
metadata:
  # The name here doesn&apos;t matter.
  name: notImportantHere
# The name of the chart to use
chartName: haproxy-ingress
# The name of the release
releaseName: hydra
# Path to a values.yaml
values: values.yaml
# Namespace to deploy to
releaseNamespace: ingress-controller
# URL to the Repo
chartRepoUrl:  https://haproxy-ingress.github.io/charts</code></pre><p>The last step is to create a <code>values.yaml</code> (or if you&apos;re happy with the defaults, you can just remove the line from <code>chartInflator.</code>)</p><p>The default behavior of the chart inflation generator is to overwrite the defaults, and keep any that aren&apos;t specified, so you can create some pretty minimal <code>values.yaml</code> files. (the one I use is below)</p><pre><code># values.yaml
controller:
  ## Annotations to be added to controller pods
  ##
  podAnnotations:
    config.linkerd.io/skip-inbound-ports: 80, 443
  extraArgs:
    watch-ingress-without-class:
  # ConfigMap to configure haproxy ingress
  config:
    proxy-protocol: &quot;no&quot;
    use-proxy-protocol: &quot;True&quot;
    config-frontend: |
      capture request header Host len 32
      capture request header X-REAL-IP len 64
      capture request header User-Agent len 200
  service:
    loadBalancerIP: &quot;10.0.7.20&quot;

  stats:
    enabled: true

  metrics:
    enabled: true

  logs:
    enabled: true

defaultBackend:
  enabled: true

</code></pre><p>And that&apos;s it! It&apos;s really simple to do, but it allows you to use Helm charts that are now managed within Kustomize.</p><p>I&apos;m currently writing some documentation for the full list of options for HelmChartInflationGenerator, which will be submitted as a PR to the Kustomize project, but it&apos;s along the same lines as above. I just wrote this because I had to look through the source code for Kustomize to find the <code>values</code> option.</p><p></p>]]></content:encoded></item><item><title><![CDATA[How I Learned to Stop Worrying and Love the Bot]]></title><description><![CDATA[An insight into running a honeypot, and sharing the statistics and insights gained with a honeypot.]]></description><link>https://blog.okami.dev/love-the-bots/</link><guid isPermaLink="false">637c0e17b30cf40001742c93</guid><category><![CDATA[security]]></category><category><![CDATA[virtual-machines]]></category><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Fri, 31 Jul 2020 23:21:26 GMT</pubDate><media:content url="https://blog.okami.dev/content/images/2020/07/kibana-1.png" medium="image"/><content:encoded><![CDATA[<img src="https://blog.okami.dev/content/images/2020/07/kibana-1.png" alt="How I Learned to Stop Worrying and Love the Bot"><p>Hello world. Lets talk about honeypots, and the helpful insights they give into the state of botting on the internet.</p><p>But first, what is a honeypot?</p><h2 id="honeypots-explained">Honeypots, explained</h2><p>A honeypot (or honeytrap); is essentially a decoy computer or container that you run to trap and track attackers. &#xA0;It can be a real machine, or a fake one designed to emulate a real one. Often honeypots are made to look appealing by appearing as though they run exploitable software. The purpose of a honeypot is to gain an insight into what attackers would do, what they would target, and what they would exploit to gain access. This better equips you to understand what methods attackers use to gain access, and enables you to better defend from future attacks.</p><p>An example of a honeypot, is the <a href="https://github.com/MalwareTech/CitrixHoneypot">Citrix Honeypot</a>. The Citrix Honeypot aims to emulate a Citrix Gateway that is vulnerable to <a href="https://support.citrix.com/article/CTX267027">CVE-2019-19781</a>. This vulnerability would allow an unauthenticated user to perform remote code execution on a Citrix Gateway (among other Citrix products). This is incredibly appealing to attackers, as if you can perform remote code execution, you can compromise firewalls, install cryptocoin miners, release malware and more. It&apos;s bad, and the Citrix vulnerability was even worse because an <em>unauthenticated</em> user could do it. That means that <em>anyone</em> could do it.</p><p>The Citrix honeypot is useful to people running Citrix systems, as they can safely<sup><a href="#ref-01">[1]</a></sup> run the honeypot without risking their real infrastructure, and with the data gathered (stuff like attacker&apos;s IP addresses), they can strengthen the security of their real infrastructure by IP banning the attackers. This might not stop them in their tracks, but it would certainly slow them down.</p><p>The majority of attacks performed on the internet aren&apos;t done by hackers in some dimly lit basement somewhere. Most of them are done by automated bots, scanning the internet for exposed ports and running scripts to break in. These kinds of attacks are frequent, and <em>annoying.</em></p><p>Honeypots, especially when used alongside monitoring and automated banning software; provide an excellent way to filter out 90% of these automated attacks; significantly reducing the attacks on your real systems.</p><!--kg-card-begin: html--><p>
    <a name="ref-01">[1] Honeypots are not always safe. Be aware that letting anyone access a system, real or not, is a risk. </a>
</p><!--kg-card-end: html--><h2 id="enter-tpot">Enter Tpot</h2><p>Because I am a nerd and I love seeing bots attempt to break in, I set up a honeypot using the open source project <a href="https://github.com/dtag-dev-sec/tpotce">Tpot</a>. Tpot is special, in that it is a preconfigured collection of honeypots listening on a number of different ports to simulate a range of different services. This casts a very wide net to catch all manner of bots.</p><p>Tpot utilizes the sandboxing capabilities of docker to make it easy to upgrade, and safer<sup><a href="#ref-01">[1]</a></sup> than running things bare metal. The best thing about tpot is that it aggregates the data from all the running services and presents the data in various forms using <a href="https://www.elastic.co/kibana">kibana</a>.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/07/kibana.png" class="kg-image" alt="How I Learned to Stop Worrying and Love the Bot" loading="lazy"><figcaption>Kibana</figcaption></figure><p>To get this running, I created a VM inside my <a href="https://www.ovirt.org/">oVirt</a> cluster and installed tpot using the ISO. I didn&apos;t want to expose my residential IP address to the internet, so instead I launched a VPS on <a href="https://blog.okami.dev/love-the-bots/Digitalocean.com/">DigitalOcean</a> running <a href="https://www.vyos.io/">VyOS</a>, and setup a wireguard link on my firewalled netork. All ports are then forwarded from the VPS to the virtual machine, exposing it directly to the internet. This gives the highest opportunity for bots and attackers to fall into my trap.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/08/tpot_network-1.svg" class="kg-image" alt="How I Learned to Stop Worrying and Love the Bot" loading="lazy"><figcaption>Diagram of components</figcaption></figure><p>I&apos;ve been running this publicly for a few weeks now and not noticed any issues so far, so lets hope for the best. Now, lets discuss the things learnt from running a public honeypot.</p><h2 id="insights-gained">Insights Gained</h2><p>Lets take a look at the statistics. Over a 15 day period, there were a total of 2,768,090 Attacks from 7,423 different IP addresses. That&apos;s Two attacks, per second, of every hour, of every day.</p><p><strong><a href="https://sec.oxide.one/kibana/app/kibana#/dashboard/Dionaea?_g=(filters:!(),refreshInterval:(pause:!t,value:0),time:(from:now-15d,to:now))&amp;_a=(description:&apos;Dionaea%20Dashboard&apos;,filters:!(),fullScreenMode:!f,options:(darkTheme:!t,useMargins:!t),query:(language:lucene,query:(query_string:(analyze_wildcard:!t,default_field:&apos;*&apos;,query:&apos;*&apos;,time_zone:Europe%2FLondon))),timeRestore:!f,title:Dionaea,viewMode:view)">Samba</a></strong></p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/07/dionaea.png" class="kg-image" alt="How I Learned to Stop Worrying and Love the Bot" loading="lazy"><figcaption>Dionaea Dashboard</figcaption></figure><p>The most popular honeypot by far, was <a href="https://github.com/DinoTools/dionaea">Dionaea</a>. This honeypot aims to emulate the popular network file sharing<sup><a href="#ref-02">[2]</a></sup> protocol <a href="https://en.wikipedia.org/wiki/Server_Message_Block">Samba (SMB)</a>, among other things. 95% of the attacks on Dionaea were targeting this protocol. </p><p>Why is the SMB protocol so popular? Well, it&apos;s tightly integrated into Windows operating systems, and has been a core component for a very long time. There are a huge number of security vulnerabilities in different versions of SMB; and it&apos;s enticing to attackers, as these vulnerabilities could easily grant an attacker with root (administrator) level access to a windows machine. </p><p>A notable example of SMB being exploited, is the famous <a href="https://en.wikipedia.org/wiki/WannaCry_ransomware_attack">WannaCry</a> series of attacks. These attacks exploited a vulnerability in the SMB protocol, which allowed an unauthenticated user to execute code on a windows machine. This code would infect the machine with the <a href="https://en.wikipedia.org/wiki/Computer_worm">worm</a>, and spread to other machines on the networks. Did I mention it <a href="https://en.wikipedia.org/wiki/Ransomware">encrypted</a> all the files on the host, demanding a ransom for the files to be decrypted? It&apos;s a trivial attack, and an easy payday for cybercriminals looking to get some bitcoin. That&apos;s probably why it&apos;s so popular.</p><p>My suggestion? Don&apos;t ever, <em>ever</em> expose SMB to the internet.</p><!--kg-card-begin: html--><p>
    <a name="ref-02">[2] SMB covers more than just filesharing; but that&apos;s beyond the scope of this article.</a>
</p><!--kg-card-end: html--><p><strong><a href="https://sec.oxide.one/kibana/app/kibana#/dashboard/Cowrie?_g=(filters:!(),refreshInterval:(pause:!t,value:0),time:(from:now-15d,to:now))&amp;_a=(description:&apos;Cowrie%20Dashboard&apos;,filters:!(),fullScreenMode:!f,options:(darkTheme:!t,useMargins:!t),query:(language:lucene,query:&apos;&apos;),timeRestore:!f,title:Cowrie,viewMode:view)">SSH &amp; Telnet</a></strong></p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/07/cowrie.png" class="kg-image" alt="How I Learned to Stop Worrying and Love the Bot" loading="lazy"><figcaption>Cowrie Dashboard</figcaption></figure><p>Next up, is Cowrie. This honeypot aims to emulate both <a href="https://en.wikipedia.org/wiki/Secure_Shell">SSH</a> and <a href="https://en.wikipedia.org/wiki/Telnet">Telnet</a>. Both of these are incredibly popular protocols, used for establishing a commandline shell over the internet.</p><p>I won&apos;t be addressing Telnet much, as it&apos;s not as popular nowadays; and it&apos;s advised to just use SSH instead<sup><a href="#ref-03">[3]</a></sup></p><p>SSH is one of the most well established tools for initiating a remote connection over the internet. It&apos;s also fairly easy to misconfigure to make it vulnerable. The problem with misconfigured SSH servers, is that if you get in; you can do a <em>lot</em>. Often you need to bruteforce passwords to get in, but once you are, game over.</p><p>The Cowrie honeypot was dealt over 290,430 attacks over the 15 day observation period. The most common attack being on the SSH implementation. <code>admin</code> and <code>root</code> were the most common usernames attempted; with 16,000 and 10,200 attacks respectively. The most common passwords were <code>1234</code>, <code>admin</code>, <code>root</code>, <code>user</code> and <code>support</code>. Basic passwords really.</p><p>What does this mean for a server operator though? </p><p>It means, that when you have SSH exposed to the internet; you need to be careful. It won&apos;t take long for bots to begin attacking you, and you <em>cannot </em>use easy to guess passwords if you care about that machine not being compromised.</p><p>Another alternative to this, and one I highly suggest, is to forgo passwords altogether and switch to <a href="https://www.digitalocean.com/community/tutorials/how-to-configure-ssh-key-based-authentication-on-a-linux-server">key based authentication</a>. This will stop 99% of attacks outright.</p><!--kg-card-begin: html--><p>
    <a name="ref-03">[3] In the majority of cases, ssh is suggested. Telnet is still used for debugging purposes due to it&apos;s simplicity.</a>
</p><!--kg-card-end: html--><p>There are a more honeypots to cover than just these two, but I don&apos;t want to make this post too long. If there&apos;s significant interest I&apos;ll do some more writeups covering some of the more unique honeypots.</p><h2 id="now-what">Now what?</h2><p>Now, I have a fairly large collection of data (50GB+ and growing). What can I do with it?</p><p>For a start I can feed the IP addresses collected from tpot, and send them to my frontline loadbalancers to drop connections from attackers. This is a work in progress, but I will be following on from this post when it is implemented.</p><p>In addition, the data collected from tpot is automatically submitted to <a href="https://www.sicherheitstacho.eu/start/main">sicherheitstacho</a>. The aim of this project is to build a realtime visualization of attacks going on around the world.</p><p>I&apos;ve also made the data public. You can take a look at the dashboards <a href="https://sec.oxide.one">here</a>. Give it a play! There&apos;s a <strong>lot</strong> that I didn&apos;t cover here.</p><p>Thanks.</p>]]></content:encoded></item><item><title><![CDATA[Losing 7 days in 4 seconds.]]></title><description><![CDATA[A journey into the mistakes I made that caused me to lose a large amount of work.]]></description><link>https://blog.okami.dev/losing-4-days-in-4-seconds/</link><guid isPermaLink="false">637c0e17b30cf40001742c90</guid><category><![CDATA[kubernetes]]></category><category><![CDATA[DevOps]]></category><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Tue, 14 Apr 2020 23:54:47 GMT</pubDate><media:content url="https://blog.okami.dev/content/images/2020/04/2-1.png" medium="image"/><content:encoded><![CDATA[<img src="https://blog.okami.dev/content/images/2020/04/2-1.png" alt="Losing 7 days in 4 seconds."><p>(Alternative title: The importance of infrastructure as code)</p><p>Hello everyone. I hope you&apos;re doing well. I&apos;m writing this blog post to share with you a recent adventure I had, and something that helped reinforce my belief that you should aim to have <em>everything</em> defined as code (where possible).</p><p>Kubernetes is the hot new thing at the moment, and I am currently migrating my infrastructure off of docker and onto this behemoth of a system for a multitude of reasons. They&apos;re the typical K8S (Kubernetes) reasons, so I won&apos;t repeat them.</p><p>One of the most important to me though was abstraction. The abstraction that is given to the end-user makes things a lot easier to deploy, manage, upgrade and more. On the downside, when you don&apos;t do abstraction <em>right,</em> it can be a major kick in the teeth when things come crashing down.</p><p>I very recently lost roughly a week&apos;s worth of work in the space of 4 seconds. This blog post is going to discuss the mistakes <em>I </em>made, and hopefully, you can learn a few things from it.</p><h2 id="the-seven-days-">The Seven days.</h2><p>To give a bit of context, I&apos;ve been running a mix of Kubernetes and Openshift clusters on my at-home server farm. My first (long term) deployment was K8s, the second was Openshift origin 4.4, and the most recent has been K8s. I call this iteration Apollo.</p><p>Apollo is special. It&apos;s been the cluster where I try to define my Kubernetes infrastructure as code and make monitoring and tracing a core feature in the cluster.</p><p>To manage the infrastructure as code (IaC), I was using a mixture of Helm and Kustomize. Helm allows for quick wins while deploying infrastructure as code, as you can deploy parameterized infrastructure, only changing the parameters as needed.</p><p>Kustomize requires you to deploy your <em>entire</em> infrastructure as code. No ifs, no buts. You want a secret, you create the YAML.</p><p>My first mistake was using a mixture of configuration management tools. It leads to fragmentation of my codebase, and as a result, maintaining quality was a little harder. When you are aiming to maintain and improve quality, try to stick to a minimal codebase. This scales with team size, so find whatever is appropriate for your team and stick with it.</p><p>Once the cluster was lovingly deployed with Ansible, I began investigating how I could monitor the state of the cluster. Having a view over the state of your environment is paramount to quality, and so any QA should push for good dashboarding and monitoring as a core component. With monitoring, you understand the health of your environment better. You recognise the weaknesses, the bottlenecks, the <em>defects.</em></p><p>Dashboards give you an understanding of the state of the system as a whole and give you visibility into where you need to test.</p><p>Within the K8S realm, Prometheus is one of the most popular monitoring tools at the moment. Combine that with a graphing tool like Grafana, and you&apos;ve got a killer combo.</p><p>I used Helm to deploy the Prometheus operator, which deployed Prometheus, Grafana and a few other components into your cluster. I saved the variables I used to deploy initially so that I could redeploy the infrastructure again, and again, and get <em>repeatable</em> results.</p><p>The Prometheus operator configures some Grafana dashboards by default, giving some great insight into the health of a Kubernetes cluster.</p><p>Once there, I logged into Grafana and began building some extra dashboards by hand. I wrote one to monitor the health of my Hypervisor cluster.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/04/unknown.png" class="kg-image" alt="Losing 7 days in 4 seconds." loading="lazy"><figcaption>Dashboard for my Ovirt cluster</figcaption></figure><p>From there, I continued building and wrote one to Monitor the heatlh of my FreeNAS system. </p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/04/2.png" class="kg-image" alt="Losing 7 days in 4 seconds." loading="lazy"><figcaption>Dashboard for my FreeNAS system</figcaption></figure><p>These manual changes provided me with even greater visibility of the health of my entire <em>system</em>, not just Kubernetes. I began to understand some of the defects within my system in a much better way.</p><p>One example of this was learning that I had a high level of queued IO requests on FreeNAS. This was causing performance issues and as a result, was causing timeouts on components reliant on FreeNAS. I had discovered a defect through the use of monitoring, was able to trace the cause back to queued IO, and fix it with creating a cache device.</p><p>Making the manual changes and creating dashboards proved to be undoubtedly useful. Although, with these manual changes lies my next mistake. <em>Manual changes. </em>I made the mistake of manually creating dashboards onto an automatic deployment. If I were to redeploy the Prometheus operator, I would lose all of these manual changes, and my dashboards.</p><p>I spent about a week creating these dashboards, working into the evening to get them configured to how I liked them.</p><p>I defined <em>most</em> of my infrastructure as code. Not <em>all</em> of it. (Foreshadowing a little bit here).</p><h2 id="four-seconds-">Four seconds.</h2><p>Four seconds is all it took to bring it all crashing down.</p><p>I had built up my dashboards to a point where I wanted to experiment with alerting me when issues are found. Stuff like CPU temperatures spiking above 80<strong>&#xB0;</strong>C for more than a few minutes. I want to be able to act quickly when an issue is found and resolve it. To do this, I wanted to configure Grafana to send me an email when an alert is triggered.</p><p>To do this, I went to my private repo, modified the helm values to set an email server and its credentials, and updated the deployment. I had assumed that the dashboards were saved into backing storage (like I would have done with Kustomize), and so updating the deployment would just load the dashboards again once it rolls out.</p><p>I pressed enter and 4 seconds later, I was notified that the deployment had completed successfully. I opened up the Grafana dashboard and was presented with a login screen.</p><p>Odd. I don&apos;t usually need to login to Grafana as I had a session token.</p><p>I logged in and noticed that the theme had changed back to the default (dark mode).</p><p>This is where my mind begins to panic.</p><p>I take a look at the dashboard list and realise they&apos;re all there; except for the ones I had created myself. The 10 minutes setting up the operator was there, but the 7 days making my dashboards were gone. S**t.</p><p>Unbeknownst to me, Helm had redeployed the <em>entire</em> operator. Meaning it wiped the slate clean and treated it like it was a completely new deployment. I had misunderstood the helm chart and assumed that it saved the configuration into storage, resulting in an iterative deployment.</p><p>This is my third mistake. Using tools I do not understand properly. Whether this is a misunderstanding of Helm, Kubernetes or something else; there is still a misunderstanding made.</p><h2 id="the-takeaway-from-all-this-">The takeaway from all this.</h2><p>So. I made a few mistakes. Let&apos;s cover them in a bit more depth so that you can recognise them in your team, and prevent them from happening.</p><p><strong>Multiple tools for the same purpose</strong></p><p>The first mistake I made was the use of multiple tools to perform the same task. While I&apos;m not saying you should use one tool for everything, you should not aim to use multiple tools that do the <em>same thing.</em> An example of this would be using Ansible and Puppet. They are both configuration management tools. Unless you are doing something highly specialised that <em>requires</em> the use of multiple tools, try to stick to one. This helps prevent knowledge of one leaking into the other and causing confusion and mistakes. In my case, I had assumed that the helm chart created storage for the Graphana dashboards (allowing them to persist) like I would have done for Kustomize.</p><p><strong>90% IaC is not acceptable.</strong></p><p>The next mistake was manual changes ontop of automated changes. I had defined 90% of my infrastructure as code, but the remaining 10% was what mattered the most.</p><p>The 10% in my case, was the custom dashboards I had created. For others, 10% could be code pipelines to deploy code. This kind of infrastructure is often easy to create, and incredibly hard to spot. Sometimes you make assumptions that the remaining 10% is not <em>that </em>important, but you won&apos;t realise that until it&apos;s <em>gone.</em> I&apos;m not going to claim I&apos;m a master at this by any stretch, but I&apos;ll try and give some tips on recognising it.</p><p>The first steps to recognising the hidden 10% are to deploy your infrastructure into a test environment with <em>only</em> the parts that are defined as code and compare it to what is <em>actually</em> deployed. This will give you a start into recognising what work needs to be done, and do not stop until you reach 100% and can <em>prove </em>it is 100%.</p><p>For some cases, 100% is not possible, and that can be ok. Mitigate against it by documenting the manual changes needed, and make sure that you can <em>prove it.</em></p><p><strong>Understand the tools you are using</strong></p><p>This links back to my first mistake somewhat. But it is important to understand the infrastructure tools you are using. I&apos;m not saying that you need to be the all-seeing guru of all things helm, but you should at least build up an understanding that allows you to do the most common tasks without needing to consult a guide. If the tool you are working with does something you did not expect, investigate why that happened and learn from it. Tools are not people, they will <em>always</em> do what you <em>tell them to do. </em>Do not sit idly and just think &quot;oh nothing will come of that&quot;, because it can.</p><h2 id="next-steps">Next steps</h2><p>Where do I go from here? I&apos;m going to drop helm. It&apos;s great for some quick and small wins, but I do not want to be in a position where I am not in control of what I am deploying, and cause an issue like this again.<br>I&apos;m probably going to wait for Openshift origin 4.4 to release (officially), and give it another go. Maybe I&apos;ll call this one Skylab. All I know is that from here I can only learn, and I hope you learnt something from this adventure too.</p>]]></content:encoded></item><item><title><![CDATA[Smashing the Stack]]></title><description><![CDATA[<p>Hello everyone.</p><p>I hope you had a good holiday season.</p><p>I took time off over the holiday to redeploy my infrastructure to make it more automated, more reliable and to be more organised.</p><p>This multi-part blog post will cover the entire journey, from planning to execution and finally, the firefighting.</p>]]></description><link>https://blog.okami.dev/smashing-the-stack-part-1/</link><guid isPermaLink="false">637c0e17b30cf40001742c8f</guid><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Sat, 25 Jan 2020 12:24:19 GMT</pubDate><media:content url="https://blog.okami.dev/content/images/2020/01/oxide.one.jpg" medium="image"/><content:encoded><![CDATA[<img src="https://blog.okami.dev/content/images/2020/01/oxide.one.jpg" alt="Smashing the Stack"><p>Hello everyone.</p><p>I hope you had a good holiday season.</p><p>I took time off over the holiday to redeploy my infrastructure to make it more automated, more reliable and to be more organised.</p><p>This multi-part blog post will cover the entire journey, from planning to execution and finally, the firefighting.</p><p>TLDR: I redeployed my infrastructure and documented it. Link <a href="https://wiki.okami.dev">here</a></p><p>Before we begin, I want to cover the reasons why I did all this.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/01/homelab_2018.jpg" class="kg-image" alt="Smashing the Stack" loading="lazy"><figcaption>My homelab in 2019</figcaption></figure><p>My infrastructure in 2019 was a single Desktop computer. Through the year it grew from that to 1 server, then to two, then to 5. </p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2020/01/DSC02296_01.jpg" class="kg-image" alt="Smashing the Stack" loading="lazy"><figcaption>my homelab now</figcaption></figure><p>While I had done my best to manage them as separate entities, my infrastructure had become large enough that I should begin treating them the way an enterprise would. Additionally, the way my infrastructure had grown had resulted in many issues arising from the gradual expansion.</p><h2 id="drift">Drift</h2><p>As my infrastructure slowly grew from one server to two, then to five; they began to drift heavily. To me, <em>drift </em>is the slight changes in configuration between machines that make them incompatible with one another when attempting to automate tasks.</p><p>One such example of drift is filesystem layout relating to Docker.</p><p>On machine A, the mount points for docker volumes might be <code>/srv/docker</code>. On machine B, this would be <code>/opt/docker</code>. While this wasn&apos;t an issue when managing 2 or 3 machines, it became an issue when I wanted to automate the deployment of docker images onto these machines using ansible. While I could have hardcoded the values to point to the correct locations, it would have resulted in much unnecessary complexity in my ansible scripts.</p><p>This kind of drift was present all over my machines, from sysctl variables to daemon configurations.</p><p>Drift happens because of a lack of proper change control and planning, and in its infancy, I didn&apos;t feel the need to enforce either.</p><h2 id="bloat-complexity">Bloat &amp; Complexity</h2><p>Bloat became a problem for my hypervisors because I wasn&apos;t treating them <em>as </em>hypervisors; I was treating them as general-purpose do all servers.</p><p>One served the purpose of a DNS server <em>and </em>a web server <em>and</em>a build server. Each of these purposes should have been segregated into a virtual machine; so that in the event that I need to shut the hypervisor down, the virtual machine could be moved onto a different hypervisor.</p><p>In hindsight, I saw it as easier to install the necessary packages and be on my way; instead of setting up a virtual machine and then the packages. This approach lacked the understanding that I would need to maintain an additional component on top of the hypervisor; rather than maintain a component <em>next to</em> the hypervisor. The system is a lot less complex when it doesn&apos;t have 20 other services running with it.</p><p>This bloat became a big problem when I had to deal with a cascading failure of one service, causing the hypervisor to enter emergency mode, resulting in the entire system being unusable until I fixed the issue. If I had put the service into a virtual machine, then I wouldn&apos;t have been in that mess.</p><p>Bloat also added to the total number of packages installed, and thus, updates also became more of an issue as the hypervisors grew in complexity.</p><p>Bloat happens when systems become too monolithic, and from a failure to segregate services properly.</p><h2 id="-a-lack-of-abstraction">(a lack of) Abstraction</h2><p>This is somewhat related to my previous two points, but it&apos;s different enough that I feel the need to point it out. In my infrastructure, I had failed to abstract my services properly. This failure to abstract meant that (for example), my NFS server was hydrogen.srv.oxide.one, <em>not nfs1.nix.oxide.one.</em></p><p>In an ideal world, I shouldn&apos;t have to care that NFS is served by hydrogen; I can point my services to nfs1, and I&apos;ll be fine.</p><p>While this isn&#x2019;t so much of an issue, it does become annoying when attempting to keep track of my infrastructure and keep things running smoothly.</p><h2 id="lack-of-repeat-ability">Lack of repeat-ability</h2><p>One of the problems that I realised I had was that my hypervisors were not repeatable. I had failed to document how to set up the machines, along with any fixes or configuration changes I had done to make that machine into what it was. Some configuration changes were useful and needed, but others were not.</p><p>What this means in practice is that if I suddenly lost the OS on the hypervisor, I would be lost when attempting to rebuild it. I had initially built it by googling my way forward, but this was not sustainable.</p><h2 id="the-goal">The goal</h2><p>With these issues in mind, I concluded that it would generally be easier to redeploy my infrastructure, rather than retrofit any fixes. I wanted this redeployment to be the <em>last</em>one I need to do. To make sure that the redeployment was the last, I need to make sure that the process was Documented and Automated, where possible.</p><p>I gave myself some goals to be able to mark this infrastructure deployment as &#x2018;successful&#x2019;. They are as follows:</p><ul><li>Be Repeatable</li><li>Do not treat any machines differently.</li><li>Keep the hypervisors as simple as possible.</li></ul><p>First, I wanted my infrastructure to be repeatable. If I wipe a hypervisor, take a fresh install of the OS and make the changes defined in the Documentation, I should end up in the same place as I was before. If not, then I need better documentation. Ensuring that this is kept, I can also eliminate drift.</p><p>Second, I should not treat my machines differently; unless there is a justifiable reason to do so. This will help to eliminate drift and complexity.</p><p>Third, keep the hypervisors as simple as possible. If something doesn&apos;t need to run on the hypervisor, don&apos;t run it. Ensure that they do not depend on each other so that they can run independently. Anything that creates a cross-machine dependency should have checks in place to ensure that cascading failures do not happen. This will eliminate my lack of abstraction and bloat.</p><p>So, I had my issues; I had my goal. Now to plan.</p><h2 id="the-plan">The plan</h2><p>I began to document everything about my machines. I had little visibility of what my infrastructure actually <em>was</em>. If I wanted to redeploy, I should at least document the mistakes I made so I don&#x2019;t make them again.This task alone took up the first week of my holiday, but it was worth it.</p><p>Having a good scope of everything in my domain gave me clarity on how much drift had been going on, both on the hardware scale and software scale. &#xA0;One of my hypervisors had over 100GB of ram, yet the other had 32. On the software scale, one of my hypervisors was setup and synchronised with LDAP; yet the other was not.</p><p>It was impressive that everything was running as smoothly as it was, in all honesty.</p><p>Once I began building out the documentation; I began to have a nagging thought in my head:</p><p>Am I creating a <em>spec, </em>or am I creating <em>documentation</em>?</p><p>A spec states what the requirements are for a system. It states how a system should function and behave, whereas documentation states how the system<em> currently </em>functions. While it was in essence Documentation, it felt a lot like a spec, and so I treated it as such.</p><p>For each component I wrote about, I determined whether it was useful to the end goal. If it was, I documented it, and it became part of the spec. If the component wasn&apos;t relevant to my end goal, and I couldn&apos;t justify the component existing, then it wouldn&apos;t become part of the spec.</p><p>One such example of this is the use of Nvidia drivers in the kernel. At some point in my infra, I had installed a GPU into each server so that it could perform hardware encoding of video. This resulted in the Nvidia kernel drivers being installed.</p><p>If I were treating this as Documentation, I would point this out and include instructions for how to set this up on a new hypervisor.</p><p>I determined that I shouldn&#x2019;t really use out of tree kernel drivers if I wanted to ensure stability; so it didn&#x2019;t become part of the spec. This kind of approach lead to me writing a hybrid Spec/Doc sheet; that would take the best parts of my old infra and build a rock-solid foundation for future expansion.</p><p>Once redeployed, the spec will become the documentation, and all will be well.</p><p>I call this infrastructure oxide one. It&#x2019;s currently at version 1, but will be updated well into the future as it is now the single source of truth for all of my machines.</p><p>The next blog post will cover how I got around to turning the spec into actions using Ansible.</p><p><a href="https://wiki.okami.dev">The wiki is available here</a></p>]]></content:encoded></item><item><title><![CDATA[The two and a half minute Kubernetes cluster]]></title><description><![CDATA[<p>(after you fulfill the prereqs first)</p><p>Hello everyone! </p><p>In this blog post I&apos;m going to be doing a write-up on creating an Ansible script to deploy a Kubernetes cluster.</p><h2 id="the-why">The Why</h2><p>I&apos;m trying to learn Kubernetes, and I tend to learn by breaking things. </p><p>I don&</p>]]></description><link>https://blog.okami.dev/2-minutes-to-kubernetes/</link><guid isPermaLink="false">637c0e17b30cf40001742c8e</guid><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Sun, 01 Dec 2019 22:42:00 GMT</pubDate><media:content url="https://blog.okami.dev/content/images/2019/12/cluster_image-1.png" medium="image"/><content:encoded><![CDATA[<img src="https://blog.okami.dev/content/images/2019/12/cluster_image-1.png" alt="The two and a half minute Kubernetes cluster"><p>(after you fulfill the prereqs first)</p><p>Hello everyone! </p><p>In this blog post I&apos;m going to be doing a write-up on creating an Ansible script to deploy a Kubernetes cluster.</p><h2 id="the-why">The Why</h2><p>I&apos;m trying to learn Kubernetes, and I tend to learn by breaking things. </p><p>I don&apos;t want to spend a significant amount of time setting up individual virtual machines, installing the packages and then setting the cluster up only to break something and tear it down again. Even with the help of the script I wrote in my <a href="https://blog.okami.dev/vm-deployments-made-easy/">previous blog post</a> it&apos;d still take a significant amount of time to get back to a usable cluster.</p><p>I wanted to build something that was able to produce repeatable builds, while still not taking out a sizeable part of my day waiting for the cluster to be ready.</p><p>There are already some fantastic projects out there that handle deploying and provisioning a cluster, such as <a href="https://github.com/kubernetes-sigs/kubespray">Kubespray</a> or <a href="https://github.com/kubernetes/kops">Kops</a>; but it feels like they&apos;re more suited for &apos;actual&apos; cloud deployments rather than my at home setup. They also add some complexity that I don&apos;t feel comfortable trying to diagnose if things go wrong.</p><p>Additionally, by automating the task I am able to boost my automation skills (which is something I love to do).</p><h2 id="the-how">The How</h2><p>The scripts are made up of two Ansible playbooks that handle different things respectively. There is the &apos;init&apos; playbook, and the deployment playbook.</p><p>&apos;init&apos; playbook involves downloading a fedora cloud image, installing the required dependencies onto that cloud image (such as Docker, Kubeadm and Container Networking plugins) and finally packaging it up into a &apos;base&apos; image. This ideally should only need to be run once and copied over to your libvirt hosts.</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2019/12/init_playbook-1.png" class="kg-image" alt="The two and a half minute Kubernetes cluster" loading="lazy"></figure><p>The deployment playbook is concerned with taking that base image, and setting up the cluster using that image as a base to build the cluster from. </p><p>I&apos;ve done it this way for a few reasons; the first being that it&apos;s a slow playbook. On each run it downloads a new fedora cloud image, installs and updates all packages, then packages it up ready for use. Overall taking about 4 minutes to complete. </p><p>The second reason is that I don&apos;t want to have to download the same packages on each node multiple times since the &apos;base&apos; image has all the packages installed already. I want to reduce the amount of unneeded repetition as much as possible to make deployment fast. I feel that downloading packages several times onto what is realistically the &apos;same&apos; machine, is unneeded.</p><p>The third reason being that doing it this way allows me to base virtual machines off of a single base image using QEMU&apos;s backing file ability. This gives the benefit of being able to save a significant amount of space, which can be in the realm of 50GB depending on cluster size.</p><h2 id="overview-of-deployment">Overview of deployment</h2><p>On the deployment side of things, the script does essentially the same things as my <a href="https://blog.okami.dev/vm-deployments-made-easy/">provisioning script</a>; but I&apos;ve rewritten it to run all commands in parallel instead of one at a time.</p><p>Once the virtual machines are ready, the script logs into the first master and runs kubeadm init; before copying the join token over and running it on each of the nodes. </p><p>Finally, it copies the kubeconf over to /tmp/kubeconf.config and applies both <a href="https://github.com/coreos/flannel">Flannel</a> and <a href="https://github.com/danderson/metallb">MetalLB</a> as those are pretty much required dependancies for an &apos;at home&apos; cluster.</p><p>And thats it! You&apos;ve got a cluster deployed and ready to go. </p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2019/12/kubectl.png" class="kg-image" alt="The two and a half minute Kubernetes cluster" loading="lazy"></figure><p>Time to break stuff!</p><h2 id="challenges-what-i-learnt">Challenges + What I learnt</h2><p>I ran into a few challenges, one of which being that Fedora 31 refused to work with both ContainerD and Cri-o; meaning I was stuck with Docker for the time being. The Cri-o issue resulted from it not accepting that fedora 31 is a <em>thing </em>and reported it as an unknown Operating System.</p><p>The other challenge I had was rewriting the script to run in parallel, as I had initially written it to run one at a time. It proved to be incredibly useful to; as it now takes less time to deploy 6 virtual machines than it <em>would</em> have been to deploy one! </p><p>Running all the scripts in parallel also taught me some best practices in expansible. I&apos;m now more confident using roles, facts and playbooks in Ansible!</p><h2 id="conclusion">Conclusion</h2><p>I hope you enjoyed reading this writeup. Source code of this playbook is available <a href="https://git.doubledash.org/okami/kubestart">here</a>, and I will be adding a README so you can run this at home :)</p><p>I almost forgot, the gif. Enjoy :p</p><!--kg-card-begin: html--><script src="https://asciinema.org/a/284712.js" id="asciicast-284712" async></script><!--kg-card-end: html-->]]></content:encoded></item><item><title><![CDATA[Virtual Machine deployments made easy]]></title><description><![CDATA[An exploration into using Ansible and Cloud-init for VM deployments.]]></description><link>https://blog.okami.dev/vm-deployments-made-easy/</link><guid isPermaLink="false">637c0e17b30cf40001742c8d</guid><category><![CDATA[libvirt]]></category><category><![CDATA[virtual-machines]]></category><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Wed, 13 Nov 2019 23:40:06 GMT</pubDate><media:content url="https://blog.okami.dev/content/images/2019/11/2019-11-13_23-43.png" medium="image"/><content:encoded><![CDATA[<img src="https://blog.okami.dev/content/images/2019/11/2019-11-13_23-43.png" alt="Virtual Machine deployments made easy"><p>Traditional Virtual Machine (VM) deployments are usually a pain to deal with on libvirt. </p><p>You open up virt-manager, click through the menus, select how many vCPUs you want, the memory, the storage and the ISO you install from.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2019/11/2019-11-13_21-08.png" class="kg-image" alt="Virtual Machine deployments made easy" loading="lazy"><figcaption>The start of a painful journey</figcaption></figure><p>Once that&apos;s done, you then need to actually install the operating system onto your VM. This is a slow process, and I imagine large cloud providers aren&apos;t secretly doing this in the background when you are spinning up instances.</p><p>Having to do all these steps manually makes deployment velocity slow, and increases <a href="https://landing.google.com/sre/sre-book/chapters/eliminating-toil/">Toil</a>. It also limits my avenues of attack when looking at ways I can automate my work. If I have to manually click through &apos;next&apos; several times before getting to a usable machine, there&apos;s not much point in me automating the rest of the stack... is there?</p><p>In this blog post, I take a look at different ways of automating the VM provisioning process, alongside eliminating a lot of the repetitive work involved with setting up a VM.</p><h2 id="part-1-approaches">Part 1: Approaches</h2><p><strong>PXEBooting</strong></p><p>PXE booting, commonly referred to as &apos;Pixie booting&apos; is a method of sending a boot image over the network to a target, and booting an image from that. Dating back to 1985, PXEbooting is well established and just about every machine with network capabilities can do it; from Rasberry Pis to Mainframes.</p><p>While PXEbooting in and of itself is not used to provision virtual machines as all it really does is send a bootloader to a machine over the network; it can be used in combination with several different tools to make it an effective method to do Virtual (and physical) machine installs.</p><p>Some of those tools are:</p><ul><li><a href="https://www.theforeman.org/">Foreman</a></li><li><a href="https://cobbler.github.io/">Cobbler</a></li><li><a href="https://maas.io/">MaaS Project</a></li><li><a href="https://fogproject.org/">FOG Project</a></li><li>Kickstart files (for RPM based distros)</li><li>Preseed files (for APT baed distros)</li></ul><p>Kickstart and Preseed files are less PXE Solutions and more predefined &apos;defaults&apos;; but they are used commonly with PXE to do unattended installs of operating systems.</p><p>Foreman, Cobbler, MaaS and FOG are PXE solutions that all offer a good experience all around; and I&apos;ve given the former 3 a try before. Foreman is useful if you&apos;re interested in managing the entire lifecycle of a machine, as it also installs puppet onto the target machine.</p><p>While they all have their respective benefits; PXE Provisioning solutions all have a common downside. Time.</p><p>When you provision a virtual machine using PXE, it has to download the base operating system from a mirror, and install all the packages. Even on the fastest of internet connections, this will take time. </p><p>When I was using foreman, it took about 10 minutes from clicking &apos;done&apos; to having a system ready for use. </p><p>How is it that Cloud services can have a machine ready for me in less than a minute? Uber fast mirrors? No. If you want a Cloud-like experience, you need a Cloud-like solution.</p><p><strong>Cloud-init</strong></p><p><a href="https://cloud-init.io/">Cloud-init</a> was born out of the idea that instead of re-downloading and installing the operating system loads of different times; why not just have one &apos;base&apos; image with the base OS installed and work from that?</p><p>What cloud-init does, is apply a set configuration changes you want to make to that <em>base </em>OS to get it to the state you want. You&apos;re always going to need to install the base operating system anyway, so why not start from when you actually make changes that make it different from any other system.</p><p>How does it work though?</p><p>On any given OS with cloud-init installed and enabled, it will check to see if an additional ISO is mounted to the Virtual Machine. In that ISO it will check for two files, user-data and meta-data.</p><figure class="kg-card kg-code-card"><pre><code class="language-yaml">#cloud-config
users:
  - name: okami
    gecos: okami
    lock-passwd: false
    sudo: ALL=(ALL) NOPASSWD:ALL
    ssh_authorized_keys:
      - ssh-rsa AAA...6eNKCKYJ8iXuHUpZJ7EOesnCR9 okami@pegasus
runcmd:
  - sed -i &apos;s/SELINUX=enforcing/SELINUX=disabled/&apos; /etc/selinux/config

timezone: Europe/London
locale: en_GB.UTF-8

write_files:
  - path: /etc/yum.repos.d/kubernetes.repo
    content: |
      [kubernetes]
      name=Kubernetes
      baseurl=https://packages.cloud.google.com/yum/repos/kubernetes-el7-x86_64
      enabled=1
      gpgcheck=1
      repo_gpgcheck=1
      gpgkey=https://packages.cloud.google.com/yum/doc/yum-key.gpg https://packages.cloud.google.com/yum/doc/rpm-package-key.gpg

  - path: /etc/sysctl.d/kubernetes.conf
    content: |
      net.bridge.bridge-nf-call-ip6tables = 1
      net.bridge.bridge-nf-call-iptables = 1
      net.ipv4.ip_forward = 1

  - path: /etc/modules-load.d/br_netfilter.conf
    content: |
      br_netfilter

power_state:
  mode: reboot
  message: Bye Bye
  timeout: 30
  condition: True

runcmd:
  - touch /etc/cloud/cloud-init.disabled  
</code></pre><figcaption>The user-data file I currently use</figcaption></figure><p>user-data is the file which tells cloud-init what you want done to the operating system. Some examples of things you can put into a user-data file are as follows:</p><ul><li>Adding users (and authorized SSH keys)</li><li>Setting the Locale</li><li>Adding Repos</li><li>Installing Packages</li><li>Updating the System</li><li>Running arbitrary commands</li></ul><p>Feel free to check out the <a href="https://cloudinit.readthedocs.io/en/latest/topics/modules.html">documentation</a> on cloud-init as there is a lot more you can do with it, than with what I just listed.</p><p>The meta-data file is where you set the machine&apos;s hostname. There are other things you can put in meta-data but for now that&apos;s all you need to know.</p><p>This allows you to create customized images, and have them ready in only a few minutes. From what I understand, this is what cloud providers use to spin up instances, as it allows you to have a very short time-to-live.</p><p>Using Cloud-init ties in very well with the use of Cloud images from Distribution vendors. A Cloud image is the &apos;base&apos; OS I mentioned earlier. Typically only a few hundred MB in size; and on startup resize themselves to cover the size of the disk you give it. These cloud-images have already got cloud-init installed and enabled so you&apos;re ready to go very quickly.</p><p>You might be thinking &quot;that&apos;s great and all, but surely I would need to generate the cloud-init ISOs manually, right?&quot;</p><p>Well, sorta? There&apos;s another weapon up my sleeve that handles all of that for you.</p><h2 id="part-2-ansible">Part 2. Ansible</h2><p>Ansible is a tool developed by Red Hat, that is essentially clever shell scripting on steroids. I use it heavily in my lab to automate processes I feel need automating, and since my adoption of Ansible; I cannot praise the tool enough.</p><p>I took an Ansible script I found on github a while back (and cannot find the source [sorry]); and modified it fairly heavily to suit my needs.</p><p>I&apos;ll give you a high level overview of what it does.</p><ol><li>Copy the base disk image from my Rsync share to a local location, renaming it {FQDN}.qcow2</li><li>Generates the hostname from one I give it on the command line and chucks it into meta-data</li><li>Generates a cloud-init ISO using the machine specific hostname, and a base user-data file</li><li>Resizes the new disk image from 200MB to 10GB</li><li>Runs virt-install, importing the new disk image and attaching the cloud-init ISO</li><li>Ejects the cloud-init ISO</li><li>Deletes the cloud-init ISO and meta-data file (no longer needed)</li></ol><!--kg-card-begin: html--><iframe src="https://gfycat.com/ifr/FilthyWideCougar" frameborder="0" scrolling="no" allowfullscreen width="640" height="470"></iframe><!--kg-card-end: html--><p>This is all done from a single command.</p><p><code>ansible-playbook -i inventory.ini image_provisioner.yml --extra-vars=&quot;runon=blackbird host_name=shuttlefish&quot;</code></p><p>After a minute or two, I can see the hostname <code>shuttlefish.vm.oxide.one</code> has been registered to my DNS server, and I can ssh into the machine.</p><p>This has dramatically increased the TTL for my virtual machines, and reduced the amount of Toil relating to VM provisioning to almost none.</p><p>This script can very easily be used in an automated setting, which I plan on doing soon (keep watch of this blog)</p><p><strong>Downsides of Cloud-init + Ansible</strong></p><p>There are downsides to taking the cloud-init approach. The biggest being that you can&apos;t use this method to provision physical machines.</p><p>If that&apos;s what you&apos;re after, you&apos;re better off going for a PXE approach. Forman is pretty damn good for this.</p><h2 id="conclusion">Conclusion</h2><p>Ansible is awesome; and when used in addition to tools like cloud-init, I am able to create Cloud-like experiences at home.</p><p>I hope you enjoyed this exploration into how I do my deployments. I mostly wrote this so I can write more blog posts that link back to this, so I hope you&apos;ll forgive if it lacks substance.</p><p>Thanks,</p><p></p><p>Okami.</p><figure class="kg-card kg-bookmark-card"><a class="kg-bookmark-container" href="https://git.doubledash.org/okami/ansible-playbooks"><div class="kg-bookmark-content"><div class="kg-bookmark-title">ansible-playbooks</div><div class="kg-bookmark-description">ansible-playbooks</div><div class="kg-bookmark-metadata"><img class="kg-bookmark-icon" src="https://git.doubledash.org/img/favicon.png" alt="Virtual Machine deployments made easy"><span class="kg-bookmark-author">GitDash</span><span class="kg-bookmark-publisher">okami</span></div></div><div class="kg-bookmark-thumbnail"><img src="https://git.doubledash.org/user/avatar/okami/-1" alt="Virtual Machine deployments made easy"></div></a></figure><p></p>]]></content:encoded></item><item><title><![CDATA[Building a new homepage]]></title><description><![CDATA[A short write-up on how I updated my site design, along with lessons learnt during the process.]]></description><link>https://blog.okami.dev/building-a-new-homepage/</link><guid isPermaLink="false">637c0e17b30cf40001742c8c</guid><category><![CDATA[Import 2022-11-21 23:47]]></category><dc:creator><![CDATA[Eve]]></dc:creator><pubDate>Fri, 25 Oct 2019 20:23:05 GMT</pubDate><media:content url="https://blog.okami.dev/content/images/2019/10/okami.dev-1.png" medium="image"/><content:encoded><![CDATA[<img src="https://blog.okami.dev/content/images/2019/10/okami.dev-1.png" alt="Building a new homepage"><p>Hello world. It&apos;s currently 08:30 on a Tuesday; and I&apos;m sat freezing on the floor in an overcrowded train to London. This is my new blog. I&apos;m trying to get better at writing about my projects and I figured a blog was a decent-ish way of doing that.</p><p>Lets begin by talking about how I rebuilt my homepage.</p><p>For me, a homepage is a place to advertise yourself. It&apos;s almost a portal into your digital identity, and although my previous homepage was alright to showcase myself; I wanted something a little more unique and interactive.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2019/10/okami.dev.old.png" class="kg-image" alt="Building a new homepage" loading="lazy"><figcaption>My old site.</figcaption></figure><p>I wanted to build a site that allowed me to really show off. One of the biggest examples of a site that <em>shows off is a </em>site called <a href="https://www.doqk.ml/">doqk.ml</a>. It&apos;s got that cool ASCII art, hacker in the 90s type feel. I guess that&apos;s what I wanted to go for with the original okami.dev, that I never quite pulled off.</p><p>I remember that the feeling I wanted to give with the original okami.dev was a feeling that you&apos;re in a terminal. That&apos;s why I went for the console font, the redhat prompt at the top and README.md type layout.</p><p>Terminals excite me. To me a terminal is one of the best tools in anyone&apos;s arsenal. I wanted to build something that almost felt like you&apos;re in one, and in a way that felt realistic. Dropping the user into a terminal would also give them a peer into who I am in a much more interactive way than just &quot;here&apos;s who I am&quot;. I could build cool little Easter eggs and custom functions to print info about me.</p><p>I figured that if I&apos;m redesigning okami.dev I should really try harder to make the user feel that they&apos;re in a terminal than I did originally.</p><p>With this in mind; I began planning out how I was going to present a &apos;terminal&apos; to the end user. There are a few approaches to doing this but in the end I went with displaying a terminal using <a href="https://github.com/xtermjs/xterm.js/">xterm.js</a> and connecting it to some form of back-end.</p><p>I initially went along with <a href="https://blog.benjojo.co.uk/post/qemu-monitor-socket-rce-vnc">benjojo&apos;s use</a> of virtual machines to build the back-end, but I I got about 4 hours into compiling Linux kernels before I realized that any edits I would make to the site would require a recompile.</p><p>I figured I could probably do it a lot easier and a lot faster with <a href="https://www.docker.com/">Docker</a>.</p><h2 id="docker">Docker</h2><p>Docker seemed like a good choice to do this kind of task, but giving people shell access to a Docker container is a bit of a security nightmare; especially given that they&apos;d be able to exploit it fairly easily.</p><p>I&apos;ve been using Docker as the workhorse for a lot of my infrastructure for quite some time. I&apos;ve enjoyed using it given how much flexibility it gives me over Virtual Machines. Automating, testing and deploying is made significantly easier.</p><p>If I wanted to go with docker I&apos;d need to put a lot of effort into hardening the image.</p><h2 id="implementation-and-security">Implementation and Security</h2><p>So, we have the front-end and the back-end. Lets build.</p><p><a href="https://github.com/yudai/gotty">GoTTY</a> ended up being what I used to connect to the Docker image as it allows you to spawn a docker container when a user connects, and eventually kills it once they disconnect.</p><p>Once I got some really basic image setup through GoTTY I shared a link with some friends to find flaws. (There were a lot).</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2019/10/myimage.gif" class="kg-image" alt="Building a new homepage" loading="lazy"></figure><p>After the first few compliments I told them to find ways to exploit this. Within a few minutes, I guess I got my wish.</p><figure class="kg-card kg-image-card kg-card-hascaption"><img src="https://blog.okami.dev/content/images/2019/10/unknown.png" class="kg-image" alt="Building a new homepage" loading="lazy"><figcaption>Don&apos;t give access to people much smarter than you.</figcaption></figure><p>Obviously, I had some work to do before releasing this.</p><p>First on the list was removing all network access. I did this before the initial &apos;alpha&apos; release as I realized that giving people network access is megadumb regardless of how much I trust them. This is done fairly easily by spawning the container with <code>-net none</code><em>.</em></p><p>The next thing I learnt was to restrict how the user interacts with the filesystem. After a few minutes of playing, someone managed to fill up the filesystem with random data. This was done with <code>dd if=/dev/urandom of=/filename.txt</code>. I fixed that by removing removing dd at build time, and mounting the container&apos;s filesystem as <code>--read-only</code>.</p><p>Third; was limiting memory and cpu limits resources. This is fairly easy to do by running the container with the following. <code>-m 32m --cpus=0.5</code>. What this does is limit the container to 32mb of ram and 0.5 of the CPU max. This prevents stuff like fork bombs as with this the container will now just crash and drop the connection if such a heinous command is entered.</p><p>Once those were out of the way, the immediate flaws were fixed. People started to get creative. </p><p>To limit how long people had to play with the container per session, I made the container timeout after 10 minutes and close. This meant people had less time to break stuff before their container would die and they&apos;d need to get a new shell.</p><p>Someone figured out you could <code>SIGKILL</code> the timeout command to drop the limit on the user, so I changed it to run as root and drop the shell to a different user.</p><p>Since then, it&apos;s been (fairly) smooth sailing. I&apos;m thankful to the people who took the time to test the image so my dumb-ass doesn&apos;t get fucked over by malicious actors.</p><h2 id="making-it-look-fancy">Making it look fancy</h2><p>After most of the bugs were fixed; I got to work with making it look cool. I wrote a few functions to print out information and skills, and built in my Zsh config to make the syntax highlighting look nicer; but otherwise it was done.</p><figure class="kg-card kg-image-card"><img src="https://blog.okami.dev/content/images/2019/10/okami.dev.png" class="kg-image" alt="Building a new homepage" loading="lazy"></figure><p>I am pretty happy with the results. I set off making something that looked like a terminal, and I feel that the end result is pretty close.</p><p>Future plans involve putting some of my code into the image so people can actually view how god awful I am at coding, but for now I&apos;m going to leave it as is.</p><p>Redesigning okami.dev taught me a lot of lessons, and it&apos;s especially taught me some great ways of securing Docker images for public use. Hopefully I&apos;ll be able to take this into future endeavors.</p><p>The code is available below:</p><figure class="kg-card kg-bookmark-card"><a class="kg-bookmark-container" href="https://github.com/okamidash/okami.dev"><div class="kg-bookmark-content"><div class="kg-bookmark-title">GitHub - okamidash/okami.dev: Code for website okami.dev</div><div class="kg-bookmark-description">Code for website okami.dev. Contribute to okamidash/okami.dev development by creating an account on GitHub.</div><div class="kg-bookmark-metadata"><img class="kg-bookmark-icon" src="https://github.com/fluidicon.png" alt="Building a new homepage"><span class="kg-bookmark-author">GitHub</span><span class="kg-bookmark-publisher">okamidash</span></div></div><div class="kg-bookmark-thumbnail"><img src="https://opengraph.githubassets.com/9cd506912d8c9d8916d636b93b213cef7e91da0f56e48637f6e606628cc2e21c/okamidash/okami.dev" alt="Building a new homepage"></div></a></figure><p>The finished site is available <a href="https://okami.dev">here</a></p><p>Thanks,</p><p></p><p>Okami.</p>]]></content:encoded></item></channel></rss>