Research · Academic & arXiv

Back to sweep

Research sweep · deep · 1948 – 2026

Information Density and Semantic Determinacy - Formal vs Natural Language

Information density, entropy and semantic determinacy in formal versus natural languages, and what it implies for specifying computation

  • Claude Opus 4.8
  • academic
  • frontier
  • blogs
  • tech

Synthesised 2026-08-27

Narrative

The entropy figures that anchor this debate are older and narrower than they are usually credited for. Shannon's 1951 guessing-game experiment put English at 0.6 to 1.3 bits per character, and Cover and King's 1978 gambling refinement tightened that to roughly 1.25 to 1.35 bits per letter. Both methods measure surprise to a predictor conditioned on prior text, not meaning, and Shannon was explicit that semantic content sat outside the theory. Hindle, Barr, Su, Gabel and Devanbu's 2012 ICSE paper reported Java n-gram cross-entropy of three to four bits, below their English baseline under the same model, and that comparison is the empirical root of the folk claim that code is denser than prose.

That claim does not survive its own follow-up literature intact. Rahman, Palani and Rigby's 2019 ICSE paper found that separators and syntax tokens, 44 percent of a Java corpus by count, drove most of the naturalness gap, and that filtering them narrows the predictability difference substantially. Casalnuovo, Sagae and Devanbu's like-with-like comparisons still found code several times more predictable than English after controlling for vocabulary, and their subsequent dual channel hypothesis, developed with Barr, Dash, Devanbu and Morgan for ICSE 2020, reframes code as carrying an algorithmic channel executed by the machine and a natural-language channel read by programmers simultaneously. The honest reading is that code is measurably more predictable, i.e. lower entropy per token, than English under matched language models, which is the opposite of the informal claim that natural language is low density and code is high density: the folk claim inverts under its own stated measure, and the real distinction operating in the literature is not bandwidth but how much of the machine's behaviour is pinned down by the surface tokens versus left to a receiver's prior.

The ambiguity-as-optimum strand supplies the missing half of that picture. Piantadosi, Tily and Gibson's 2012 Cognition paper argues formally that efficient communication systems are ambiguous whenever context carries information, tested against English, German and Dutch corpora. Levy and Jaeger's Uniform Information Density work, building on Aylett and Turk's earlier prosodic finding that speakers shorten redundant material, and Gibson, Piantadosi, Brink, Bergen, Lim and Saxe's 2013 noisy-channel account of word-order variation together model comprehension as Bayesian decoding of a corrupted signal against a shared prior. Coupé et al.'s 2019 Science Advances estimate of roughly 39 bits per second across 17 languages is frequently cited as proof of a universal information rate, but it measures syllables against speech rate, not semantic content, and is easy to over-read as settling a question it does not address.

On specification itself, Dijkstra's 1978 EWD667, Brooks' 1987 essential-accidental distinction and Naur's 1985 theory-building argument converge on a claim later empirical work has only partly tested: that formality removes context-dependence but does not by itself remove underdetermination, because the theory of a system lives in its builders rather than its artefact. Industrial formal-methods evidence is genuinely mixed. Newcombe et al.'s 2015 AWS report of TLA+ catching design bugs in DynamoDB and S3 sits against a 2024 systematic review of a decade of industrial TLA+ practice that documents persistent adoption barriers around learning curve and tool integration. The LLM-era literature (METR's HCAST and RE-Bench, Code Roulette, and the 2025 study on prompt underspecification finding LLMs silently infer unstated requirements 41 percent of the time but regress twice as often when they do) treats the same underdetermination as a measurable quantity conditional on a prompt, with property-based testing and formal verification proposed, with mixed early results, as the residual disambiguating layer.


Sources

ID Title Outlet Date Significance
a1 A Mathematical Theory of Communication Bell System Technical Journal 1948 Establishes Shannon entropy as a measure of average surprise given a probability distribution, the founding definition that later informal appeals to information density borrow without their statistical apparatus.
a2 Prediction and Entropy of Printed English Bell System Technical Journal 1951 Shannon's guessing-game experiment yields the canonical estimate of 0.6 to 1.3 bits per letter for English, obtained from human next-symbol prediction rather than corpus statistics, and explicitly brackets meaning outside the calculation.
a3 A Convergent Gambling Estimate of the Entropy of English IEEE Transactions on Information Theory 1978 Cover and King's sequential-betting method refines Shannon's estimate to about 1.25 to 1.35 bits per character, the most cited empirical entropy figure for English used as a baseline in later code-versus-language comparisons.
a4 On the Naturalness of Software ICSE 2012 (Proc. 34th Intl. Conf. Software Engineering) 2012 Hindle, Barr, Su, Gabel and Devanbu report n-gram cross-entropy of Java corpora between three and four bits, lower than English cross-entropy under the same model, the paper that seeded the folk claim that code is more information-dense.
a5 Natural Software Revisited ICSE 2019 2019 Rahman, Palani and Rigby show separators and syntax tokens account for 44 percent of Java tokens and drive most of the measured naturalness effect, so once punctuation and stopwords are filtered on both sides the code-versus-English predictability gap shrinks sharply.
a6 Studying the Difference Between Natural and Programming Language Corpora Empirical Software Engineering (Springer) / arXiv 2018 Casalnuovo, Sagae and Devanbu attempt a like-with-like comparison across languages and models and find code four to five times more predictable than English even after controlling for vocabulary size, tightening but not overturning the Hindle result.
a7 A Theory of Dual Channel Constraints ICSE 2020 New Ideas and Emerging Results 2020 Casalnuovo, Barr, Dash, Devanbu and Morgan propose that code simultaneously carries an algorithmic channel executed by the machine and a natural-language channel (identifiers, comments, style) read by humans, giving a formal handle on why source code is neither pure formalism nor pure prose.
a8 The Communicative Function of Ambiguity in Language Cognition 2012-03 Piantadosi, Tily and Gibson give an information-theoretic argument, tested on English, German and Dutch corpora, that efficient communication systems are ambiguous whenever context is informative, reframing ambiguity as an optimum rather than a defect.
a9 Different Languages, Similar Encoding Efficiency: Comparable Information Rates Across the Human Communicative Niche Science Advances 2019-09 Coupé, Oh, Dediu and Pellegrino measure information per syllable against speech rate across 17 languages and report convergence near 39 bits per second, a figure frequently detached from its methodology in popular retellings.
a10 Speakers Optimize Information Density Through Syntactic Reduction NeurIPS 19 (NIPS 2006 proceedings) 2007 Levy and Jaeger's original Uniform Information Density paper shows speakers selectively retain optional markers such as complementizer 'that' precisely when omitting them would create a local density spike, founding the UID research programme.
a11 The Smooth Signal Redundancy Hypothesis: A Functional Explanation for Relationships Between Redundancy, Prosodic Prominence, and Duration in Spontaneous Speech Language and Speech 2004 Aylett and Turk's prosodic-duration study is the direct precursor to Uniform Information Density, showing redundant material is spoken faster and less prominently, evidence that acoustic-level redundancy already tracks predictability before Levy and Jaeger formalised the claim for syntax.
a12 A Noisy-Channel Account of Crosslinguistic Word-Order Variation Psychological Science 2013 Gibson, Piantadosi, Brink, Bergen, Lim and Saxe model comprehenders as Bayesian decoders integrating a noisy signal with a prior over likely meanings, the theoretical basis for treating ambiguity as recoverable rather than costly, conditional on a shared prior.
a13 On the Expressive Power of Programming Languages Science of Computer Programming 1991 Felleisen supplies the first formal definition of relative expressiveness via macro-translatability, giving a rigorous alternative to informal claims that one notation is 'denser' or 'more powerful' than another.
a14 Notation as a Tool of Thought Communications of the ACM (1979 ACM Turing Award Lecture) 1980 Iverson's founding argument for notational density as cognitive leverage, the strongest primary-source statement of the position that terse formal notation extends reasoning rather than merely compressing it.
a15 On the Foolishness of "Natural Language Programming" (EWD 667) E. W. Dijkstra Archive, University of Texas at Austin 1978 Dijkstra's original argument that natural language's unavoidable ambiguity, not its verbosity, disqualifies it for programming a machine with no shared context to disambiguate against, the canonical anti-NL-programming text.
a16 No Silver Bullet: Essence and Accidents of Software Engineering IEEE Computer, vol. 20 no. 4 1987 Brooks separates essential complexity (inherent in the problem) from accidental complexity (introduced by tooling and notation) and argues a sufficiently precise specification of the essence already is the program, bounding what any notation, formal or natural, can remove.
a17 Programming as Theory Building Microprocessing and Microprogramming 1985 Naur argues the working theory of a program lives in the minds of its builders and is never fully recoverable from source code or documentation alone, an argument that cuts against any claim that a sufficiently formal artefact could substitute for shared understanding.
a18 A Tutorial Introduction to the Minimum Description Length Principle arXiv (Grunwald, later published in Advances in Minimum Description Length) 2004 Lays out Rissanen's MDL as a computable approximation to Kolmogorov complexity and states the invariance theorem's fine print (agreement only up to an additive constant, and uncomputability of the true minimum), the formal limits on using shortest-description length as a practical density measure.
a19 Detecting Ambiguities in Requirements Documents Using Inspections 1st Workshop on Inspection in Software Engineering (WISE) 2001 Kamsties, Berry and Paech distinguish linguistic ambiguity from requirements-engineering-specific ambiguity and report inspection-based detection results, part of the empirical base for claims about defect rates traceable to specification ambiguity.
a20 How Amazon Web Services Uses Formal Methods Communications of the ACM (Newcombe, Rath, Zhang, Munteanu, Brooker, Deardeuff) 2015 The most-cited industrial success report for TLA+, describing bugs found in DynamoDB, S3 and EBS designs before implementation, used as the benchmark case for what formal specification demonstrably delivers at scale.
a21 A Systematic Literature Review on a Decade of Industrial TLA+ Practice arXiv 2024-11 Surveys a decade of published TLA+ case studies and names the recurring adoption barriers, steep learning curve, abstraction selection, and poor tool integration, giving an evidence base for why formal methods stayed a minority practice despite documented successes.
a22 Formal Methods in Dependable Systems Engineering: A Survey of Professionals from Europe and North America arXiv (empirical software engineering survey) 2018-12 Practitioner survey data on why formal methods are used or avoided in dependable-systems work, complementing the TLA+ literature review with cross-tool, cross-organisation evidence on adoption drivers.
a23 HCAST: Human-Calibrated Autonomy Software Tasks METR / arXiv 2025-03 METR's 189-task benchmark grounds AI agent software-engineering performance against human completion-time baselines, the methodological anchor for measuring how far natural-language task specifications actually constrain agent behaviour.
a24 RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents Against Human Experts METR / arXiv 2024-11 Reports frontier agents reaching average human-expert scores on AI R&D engineering environments under an 8-hour compute budget, evidence relevant to how much residual behavioural variance persists once a task specification is fixed.
a25 Code Roulette: How Prompt Variability Affects LLM Code Generation arXiv 2025-06 Directly measures how meaning-preserving rewordings of the same coding prompt shift generated-code correctness and structure, empirical evidence for prompt underdetermination of LLM behaviour rather than assertion.

We use analytics cookies to understand site usage and improve the service. We do not use marketing cookies.