← Portfolio home

Technical article · Markdown · LLM output

When LLM Output Breaks Nested Markdown Code Blocks, and How Indented Blocks + Bookmarklets Stabilize Them

A practical study of delimiter collisions, parser differences, and a more stable pattern for copying nested code examples between common Markdown environments.

Abstract / Overview

Markdown generated by language models can break when code blocks are nested. This article compares the Markdown specification with observed behavior in chat interfaces and editors, then proposes a practical pattern designed to survive copying and reuse. The goal is operational stability, not a universal or theoretically perfect solution.

1. Introduction: What breaks, and why it matters

When asked to explain a topic in Markdown or put an answer in a code block, a model may use the same fence for both the outer document and an inner example. A parser then reads the inner closing fence as the end of the outer block. The chat interface may appear correct even when copied plain text is not.

This discussion is based mainly on ChatGPT experiments, with behavior that can also occur in other LLM workflows. It does not claim that every renderer behaves identically.

2. Markdown code blocks: a minimal specification review

2.1 Indented code blocks

Original Markdown supports code blocks by indenting each line with four spaces (or a tab). This older form is widely supported and naturally avoids fence collisions when nested. It is less convenient to read and edit, and it does not provide a standard language label for syntax highlighting.

Indented code block structureEach nested line is shifted right by four spaces inside the outer code block.Outer Markdown block····let sum = 10 + 20;····console.log(sum);
Figure 1. Conceptual view of a nested code sample represented with four-space indentation.

2.2 Fenced code blocks

CommonMark and GitHub Flavored Markdown define fenced blocks using at least three matching backticks or tildes. A language identifier may follow the opening fence. Fence lengths and delimiter choice matter when examples contain other fences.

Fenced code block structureAn opening fence includes a language tag, followed by code and a matching closing fence.```javascriptconst answer = 42;```
Figure 2. A fenced block has an opening delimiter, optional info string, body, and matching closing delimiter.

3. Why nested blocks break: specification versus practice

The common failure is delimiter collision: an outer block and an inner example both use three backticks. The parser may treat the inner delimiter as closing the outer block, so following paragraphs become ordinary text or are rendered as code. Language tags and post-processing can make the visible result harder to predict.

Fence collisionMatching triple-backtick delimiters collide, causing the outer example to close early.```markdown outer opens ```javascript inner opens ``` parser may close outer here
Figure 3. Identical delimiter sequences can make the intended nesting ambiguous to a renderer.

Markdown has several related dialects, including CommonMark, GFM, Pandoc, and editor-specific renderers. Edge cases involving nesting, blockquotes, inline code, and post-processors can differ. Content that looks right in one interface may not survive copy and paste into another.

4. LLM-specific behavior

In the experiments, models strongly favored triple backticks, appeared less reliable at maintaining an outer fence across long output, and sometimes produced different results in rendered UI and copied plain text. Training patterns and service-side formatting may both contribute. The specification alone does not explain all observed failures.

Practical implication: inspect the copied text in its destination renderer when the structure matters. A chat preview is not a substitute for that check.
Copy and paste can expose Markdown structure errorsLLM text is rendered in a chat interface, copied as text, then parsed again by an editor where a fence collision may become visible.LLM outputMarkdown textChat UIrendered viewCopied textplain textEditor parserbreakage may show
Figure 4. Rendering, copying, and parsing are separate stages; a mismatch may surface only after pasting.

5. Experiments: six patterns

Six nested-code patterns were compared in ChatGPT, then selected outputs were copied into Obsidian. The first three did not produce usable results in the observed runs. Cases four through six were taken forward; cases four and five remained the most practical candidates.

1 · Same triple-backtick fence

Markdown outside, JavaScript inside. Delimiters collide.

2 · Short outer, longer inner fence

Three backticks outside and four inside. Length conventions can vary in practice.

3 · Bash outer, JavaScript inner

Language labels do not prevent delimiter collision.

4 · Backticks outside, tildes inside

Separate delimiter characters avoid direct collision.

5 · Indented inner code

Four-space indentation avoids an inner fence and is broadly compatible.

6 · Indented hybrid

Indent an inner fenced example within the outer Markdown block.

To display a calculation in JavaScript:

    let sum = 10 + 20;
    console.log("The total is " + sum + ".");

The inner sample is indented instead of opened with another fence.

Cases 4–6 were pasted into Obsidian for a second check. The original study also recorded ChatGPT renderings and paste results; this clean static edition summarizes those observations without embedding service-interface screenshots.

6. Result: a rule based on indented style

Rule for nested code blocks
Inside a fenced code block, represent inner code using four-space indentation. Do not put a fenced code block inside another fenced code block.

This is a practical rule, not a guarantee. Some environments may need adjustment. The original workflow used a bookmarklet to reinsert the instruction when preparing a prompt; applying it only when needed avoids imposing it on every response.

7. Why indentation was selected

Tildes can be a useful alternative because they differ from the backticks commonly used by models and can retain language identifiers. However, in the observed workflow, getting a model to consistently preserve the alternate fence was less reliable. A later processing step may also normalize delimiters.

Indented blocks are widely supported and less exposed to fence collisions or normalization. Their trade-offs are reduced readability in raw text and no standard language annotation. In controlled environments, tildes remain a reasonable option to test.

8. Bookmarklet and practical limits

A browser bookmarklet can prepend the indentation rule to the current prompt before requesting Markdown output. This allows the instruction to be applied on demand. The exact interaction depends on the chat interface, which may change; browser extensions could implement a similar workflow.

9. Summary

References

  1. CommonMark Project, CommonMark Specification 0.31.2.
  2. GitHub, GitHub Flavored Markdown Spec.
  3. John Gruber, Markdown: Syntax.