使用 Gleam、Org-Mode 和 Pandoc 撰写博客
Blogging with Gleam, Org-Mode and Pandoc

原始链接: https://byzantine-systems.github.io/blogging-with-gleam-org-mode-and-pandoc/

本文介绍了一套以 Emacs 和 Org-mode 为核心的静态站点发布流水线。Org 文件是唯一的事实来源,支持可执行源代码块、BibLaTeX 引用和可复现的 d2 图表。该方案不要求站点生成器直接解析 Org,而是由 Gleam 程序调用 Pandoc,将文章转换为 GitHub 风格的 Markdown,并使用小型 Lua 过滤器处理诗歌、参考文献和图片路径。 随后,Gleam、Blogatto 和 Lustre 将 Markdown 转换为经过类型检查的 HTML、RSS 和静态资源。Nix flakes 与 devenv 用于锁定依赖并提供可复现的开发环境;`make run` 和 `make dev` 负责完成整个构建过程。 尽管对于一个简单站点来说,这套技术栈似乎显得过于庞大,但每个工具都承担着明确而专注的角色。作者将这种“有益的臃肿”视为一种连贯的替代方案,可以省去在多个专用编辑器和应用之间切换的麻烦。整套系统的核心仍然是 Emacs:Org-mode 在一个交互式环境中整合了文章写作、代码、计算、结构组织、引用、图表和发布等功能,提供了纯 Markdown 难以自然复现的能力。

一篇题为《使用 Gleam、Org Mode 和 Pandoc 写博客》的 Hacker News 帖子引发了关于博客写作复杂度的短暂讨论。一位评论者偏爱 WordPress,因为这样他们可以专注于写作,而不必处理 Markdown、持续集成(CI)流程或自定义标记;不过,他们也承认 WordPress 存在安全弱点,而且 Ghost 的编辑器功能有限。另一位评论者将这场争论形容为一张经典的“钟形曲线”梗图:人们可能会把博客工具搞得过于复杂,而最简单的做法就是不管使用什么平台,直接开始写作。抓取该页面时,这篇帖子获得了 14 分和两条评论。
相关文章

原文

I’ve stumbled across a method of composing programs that excites me very much. In fact, my enthusiasm is so great that I must warn the reader to discount much of what I shall say as the ravings of a fanatic who thinks he has just seen a great light.

(Knuth 1984, 1)

Welcome to the first Byzantine Systems blog post. This is not my first attempt at building a website around Org-mode. I have already written about the road to Emacs maximalism in my own personal blog, where Emacs, a few questionable scripts and Nix held everything together. This time I wanted to keep the part that worked (writing everything in Emacs) while making the publishing side smaller, typed and easier to understand.

The end experiment uses the following tools:

  • Orgmode as the source format and Emacs as the writing environment.
  • d2 for making declarative diagrams.
  • Babel to make org even more powerful.
  • Pandoc as a robust, mature converter.
  • Nix, flakes and devenv to pin the build environment.
  • Gleam and Blogatto for a small, typed static-site pipeline.
  • Lustre for defining the surrounding HTML without introducing a separate template language.

The last two points look too overkill for a simple static page, but I'll let you jugde by way we create the page that lists all posts. It's just a Gleam function returning a Lustre element:

The post index is ordinary, type-checked Gleam code.

pub fn posts_page(posts: List(Post(msg))) -> Element(msg) {
  let sorted = list.sort(posts, fn(a, b) { timestamp.compare(b.date, a.date) })

  layout("Posts", [
    html.h2([], [html.text("Posts")]),
    case sorted {
      [] -> html.p([], [html.text("No posts yet.")])
      _ -> html.ul([], list.map(sorted, post_link))
    },
  ])
}

The tl;dr version of this blogpost looks like this:

One important detail, that previous diagram was generated inside Emacs, I want to keep the idea of "diagrams as code" alive, and for now d2 has been serving me well. This is what the diagram code looks like without the Emacs magic to keep only the result (the image).


direction: right

style.fill: "#ffffff"

# Default node styling
*.shape: rectangle
*.style: {
  border-radius: 8
  stroke-width: 2
  font-size: 12
}

# Default edge styling
(* -> *)[*].style: {
  stroke: "#64748b"
  font-color: "#475569"
  stroke-width: 2
  font-size: 10
}

emacs: "1. Emacs + Orgmode\nWrite raw content (.org)" {
  style: {
    fill: "#f3e8ff"
    stroke: "#a855f7"
    font-color: "#581c87"
  }
}

pandoc: "2. Pandoc + plugins\nConvert Org to Markdown (.md)" {
  style: {
    fill: "#e0f2fe"
    stroke: "#0284c7"
    font-color: "#075985"
  }
}

blogatto: "3. Blogatto\nGenerate final HTML" {
  style: {
    fill: "#dcfce7"
    stroke: "#16a34a"
    font-color: "#14532d"
  }
}

emacs -> pandoc: Export
pandoc -> blogatto: Build

Moderately related, but my close friend Marcos Magueta made an interesting video called GNU Shepherd is the abstraction for a better computing age, there he argues about BLOAT, a word so despised in the UNIX world, but most people do so without:

  • Properly defining what BLOAT really is.
  • Without knowing waht GOOD BLOAT can give you.

Emacs, when taken from a UNIX philosophy perspective, is pure bloat. Indeed, Emacs does not follow UNIX's "do one thing well". One can even argue that is is actually good, the lack of unified environment ends up causing userland BLOAT anyway. By this very definition, Emacs would be good bloat.

Consistency requires a concentrated effort on the part of some central body that promulgates standards. Applications on the Macintosh are consistent because they follow a guidebook published by Apple. No such body has ever existed for Unix utilities. As a result, some utilities take their options preceded by a dash, some don’t. Some read standard input, some don’t. Some write standard output, some don’t. Some create files world writable, some don’t. Some report errors, some don’t. Some put a space between an option and a filename, some don’t.

(Garfinkle, Weise, and Strassmann 1994, 25–26 chap.2)

Also, leaving Emacs would be admiting defeat. Writing a blogpost (or notes) for me is more than just writing paragraphs. I like to put source code around, I sometimes like to export my notes to , I like having proper bibtex citations working at all times. If every one of those requires another application, the work becomes a tour of windows (or TUIs) which happen to contain fragments of the same idea. The long term goal is eventually moving everything inside Emacs, but that day is not today.

(…) trivial usage of Org-mode is nothing more than text editing, from which point the user can start to add special plain text Org-mode elements to the document. Org-mode is therefore easy to adopt and aims to be a general solution for authoring projects with mixed computational and natural languages. It supports multiple programming languages, export targets, and work flows.

(Schulte et al. 2012, 2)

We already have Markdown, right? It is ubiquitous, readable and supported by almost every static-site generator. Indeed, this site uses the Markdown pandoc creates it and blogatto consumes it. I just do not write it directly.

The reason being that Orgmode is a far superior format, whose only downside is that you have to be inside Emacs to fully enjoy (and that seems like an okayish tradeoff). If you really want it, .org can stay as a simple lightweight markup format… or you can run arbitrary source blocks inside your org document, with custom command arguments and even a place for the generated results.

That's what makes it really powerful, that's why I like to embed d2 diagrams inside my org notes in such a way that I can always re-generate images/diagrams on demmand. It's a small reproducible (research-like) environment that is very similar to the famous Jupyter Notebooks.

Org-mode extends Emacs with a simple and powerful markup language that turns it into a language for creating, parsing, and interacting with hierarchically-organized text documents. Its rich feature set includes text structuring, project management, and a publishing system that can export to a variety of formats. Source code and data are located in active blocks, distinct from text sections, where "active" here means that code and data blocks can be evaluated to return their contents or their computational results. The results of code block evaluation can be written to a named data block in the document, where it can be referred to by other code blocks, any one of which can be written in a different computing language. In this way, an Org-mode buffer becomes a place where different computer languages communicate with one another. Like Emacs, Org-mode is extensible: support for new languages can be added by the user in a modular fashion through the definition of a small number of Emacs Lisp functions.

(Schulte et al. 2012, 7)

Markdown can imitate pieces of this through extensions and external tools, but then the authoring model is the accidental intersection of a particular parser, editor and plug-in collection. Org-mode's features belong to the same coherent system. More importantly, they are interactive in the editor instead of being instructions which only become meaningful after a build.

One tempting approach would be to teach blogatto how to parse Org directly. I'll make sure to explore this in the long run and even push to the main repo. For now, pandoc already understands Org, already emits GitHub Flavored Markdown and already has a citation processor, so the conversion stage can stay boring:

pandoc -s org/posts/example.org \
  -t gfm \
  --citeproc \
  --bibliography=priv/bibtex/emacs.bib \
  -o blog/posts/example/index.md

The real command is assembled by src/convert.gleam. It finds every Org page and post, adds all the BibLaTeX databases under priv/bibtex and writes the Markdown into the directory structure Blogatto expects. Converted files are build products and are ignored by Git. The org/ directory is the only source of truth.

There are three small Lua filters for the cases where a generic conversion is not quite enough:

  • verse.lua preserves Org verse blocks as structured HTML;
  • bibliography.lua keeps the generated references correctly grouped; and
  • images.lua translates file paths which make sense beside an Org document into paths which make sense on the published website.

The converter is a normal Gleam program: it creates output directories, skips files which are already up to date, invokes Pandoc and reports failures as Result values. Once conversion finishes, Blogatto reads the generated posts, builds their pages and RSS feed, and copies static assets into dist/. Lustre supplies the HTML elements used by the site layout.

There is deliberately not much more to say. I could be using Org-Publish to automate most of this, but there are some quality of life features that I always find it hard to make them work when generating html with it (syntax highlighting is just one of them).

Nothing really changed here, Nix is still the ducktape that makes everything work. Entering nix develop provides gleam, pandoc, make and the other build-time dependencies. Devenv makes that shell pleasant to use interactively, while the same flake reproduces the build environment without relying on whatever happens to be installed globally.


# ...
packages = {
  default = inputs'.nix-gleam.packages.buildGleamApplication {
    src = ./.;
    # `fs` (transitive dependency from blogatto -> filespy) builds a
    # NIF using the `pc` rebar3 plugin, bundle it so pure builds don't
    # hit hexpm.
    rebar3Package = pkgs.rebar3WithPlugins {
      plugins = with pkgs.beamPackages; [ pc ];
    };
  };
};

# nix fmt + nix flake check (auto-wired by flakeModule)
treefmt = {
  projectRootFile = "flake.nix";
  programs.gleam.enable = true;
  programs.nixfmt.enable = true;
};

devenv.shells.default = {
  # ...
  packages =
    with pkgs;
    [
      gnumake
      pandoc
    ]
    ++ lib.optionals stdenv.isLinux [ inotify-tools ]
    ++ [ config.packages.default ];

  languages.gleam = {
    enable = true;
  };

  languages.lua = {
    enable = true;
    lsp.enable = true;
  };

  enterShell = ''
    echo "Starting Development Environment..."
  '';
};

# ...

From inside the development environment the complete build is intentionally mundane:


make run
# or
make dev

This website has a (somewhat absurd) number of tools for a simple static page. Each one of these tools, however, has a narrow job. Org-mode holds the rich source document. Pandoc reduces it to a portable interchange format. Blogatto turns that format into a site. Gleam + Lustre give me a type-checked project that generates all the static html. Nix gives me a reproducible development environment.

The centre of the system is still Emacs, it's my (stubborn) decision to keep the act of writing, the structure of the document and the computations which support it all in one place. Org-mode is superior to Markdown for that job because it is a document model I can interact with while thinking.

Garfinkle, Simson, Daniel Weise, and Steven Strassmann. 1994. UNIX-Hater Handbook. IDG Books Worldwide, Inc.

Knuth, Donald Ervin. 1984. “Literate Programming.” The Computer Journal 27 (2): 97–111.

Schulte, Eric, Dan Davison, Thomas Dye, and Carsten Dominik. 2012. “A Multi-Language Computing Environment for Literate Programming and Reproducible Research.” Journal of Statistical Software 46: 1–24.

联系我们 contact @ memedata.com