======================================================================================== Fluid Metadata: The Future of Frictionless Writing without Strict Rules nor Rigid Syntax ======================================================================================== :Author: https://mememodoki.jp :License: CC BY 4.0 :Note: Compiled and synthesized by the author with Google AI assistance :Summary: A rationale-driven blueprint for building low-cost, human-first data lakes that bypass the quadratic attention tax and the llms-full.txt maintenance curse. :Project: Decentralized AI-Native Web Architecture test site, https://mememodoki.jp/ :Context: Rejecting rigid IT standards in favor of fluid, natural language text hooks. :Status: Grassroots Proposal / Theoretical Blueprint :Related: low-cost-information-sharing-in-AI-era The English part below, which is generated by AI, is not my Japanese part translation. I wrote some in Japanese parallel to the session. This file is a merged one. 英語版と日本語版と記載しているが、**翻訳として対応しているわけではない**。 まずは Google AI 君とのセッションで提供された英語版を掲載する。 もともと日本語版ですらきちんと文書化することに手を焼いていた。 英語で発信するほうがよりよいが、筆者の英語力では更に厄介と思っていた。 そんなことを考えつつ AI 君とタイトルに関して対話を続けていたときに、この英語版が生じたのだ。 ライセンス的にも同梱してよいらしく、Google の寛容さに甘んじさせて頂く。 正直に言って、筆者が執筆していた日本語版 (このファイルの後半に記載している) より、 英語版の AI 出力テキストのほうが簡明で的確である。脱帽である。 .. START_ENGLISH_VERSION_BY_AI The Ingestion Dilemma ===================== Modern Large Language Models suffer from a documented structural limitation known as the "Lost in the Middle" phenomenon. When forced to process massive, multi-megabyte files, the model's finite attention window disproportionately focuses on information at the absolute boundaries, letting the center fade into digital noise. To solve this, current industrial trends attempt to enforce rigid parallel mirrors of the web—specifically the `llms-full.txt` standard. This structure passes a heavy, unsustainable maintenance tax onto human creators, forcing them to format, compile, and flatten their entire documentation trees into singular text dumps. We reject this curse. The burden of chunking, parsing, and context routing belongs entirely to the giant tech platforms, running on their own operational GPU budgets. The Mathematics are Already Solved ================================== #Context SemanticResiliency #Depends CoreVectorMathematics The anxiety that humanity must invent an unforgiving, strict syntax for metadata is a hangover from legacy computer science. Modern AI engines do not rely on line-by-line clockwork code interpreters. Through vector embeddings and semantic token weights, the underlying math groups conceptual intents together automatically. Whether a creator formats a prerequisite boundary as a top-of-file header list (:Depends: TopicName), an inline social hashtag (#Require TopicName), or a casual plain-text sentence, the AI's self-attention mechanism identifies the underlying semantic anchor effortlessly. Rigid strictness is no longer a requirement for long-term textual durability. The Four Foundational Hooks =========================== To maximize writing efficiency and protect human "laziness," we propose a minimal, non-strict vocabulary guide. These anchors can be dropped anywhere in a continuous text stream, using whatever typographic wrapper fits the writer's immediate creative flow: 1. **Scope / Part**: Conversationally defines where a text block sits inside a long narrative. 2. **Depends / Require**: Signals conceptual prerequisites so the AI can cross-reference its data lake. 3. **Refer / Source**: Establishes origin and credit, allowing future chat UIs to restructure information with appropriate legal citations rather than scraping anonymously. 4. **Context / Next**: Indicates immediate conceptual trajectory, keeping the narrative thread intact even when crawlers aggressively slice the document. The Universal Win-Win ===================== #Refer CC-BY-4.0-Data-Lake-Initiative By abandoning the requirement to be "pretty printable on A4 paper" or strictly parseable for brittle compilers, we enter a cooperative era between human authors and AI systems. Human writers retain total freedom of layout and low-friction creativity. AI platforms receive higher-quality, context-aware information directly from the source. Future users gain access to deeply nuanced, trustworthy answers that they can easily double-check and verify. Even privacy-respecting ad systems win by targeting context rather than tracking individual cookies. We keep our layout unified, automated, and grassroots. We let the machines adapt to us, not the other way around. .. END_ENGLISH_VERSION_BY_AI .. from here what author attempted to write in Japanese. これ以降は著者が AI 君との対話と並行して書いていた部分となる。 柔軟なメタデータと最小限のルールで低コストな情報源を模索する =============================================================== AI による学習とその推論能力を前提に、開かれた次世代情報基盤を柔軟で低コストな手法で模索する。 狙いは「技術的な対応で風化してしまわない情報源の構築と低コストでの維持」であって、 今現在存在する特定の問題を解決するための局所的技術階ではない。 他方で Lost in the Middle などの問題に対してファイル分割や llms-full.txt といった、 今現在の AI の抱える問題への「局所的な技術解」が標準として採用されようとしている。 一部の都合で全員に対して本質的に不要なコストが転嫁されかねない。 AI を軸に自然言語の文脈解析を前提にできる世になりつつあり、 それが今後進んでいく前提に立つなら、 必要な標準基準とはそれ自体が極めて小さく、 全員にとって低コストなものであるべきだ。 Lost in the Middle に固執した技術都合を避ける ------------------------------------------------------- 典型的な技術的解は llms-full.txt やファイル分割になる。 しかしそれは技術屋視点の、しかも今現在のツールとリソース都合に基づいたものだ。 llms-full.txt は執筆者サイドへのコスト転嫁をしている。 タイトルの変更などで llms-full.txt という便宜上のファイルの内容も、 適時更新してアップロードしなおさないとならない。 それは完全に不要な技術面対応とコストを執筆者側、 ひいては将来のあらゆる著作者に対して課すことになる。 単純なサイト構造を案内するだけの llms.txt と異なり、到底容認できない。 長大なテキストの内部メタデータはそのテキストファイル内にあるべきなのだ。 しかもそのあり方も厳密な古典的構文解析機械の便宜ではなく、 執筆している側、そのファイルのまま (偶然) 読んでいる読者の側、 つまりは人間側のコストを意識しておくべきだ。 結論: 厳密な構文もルールも定めない。 厳密な構文規則も配置ルールも無し: ヒトが読んで把握できれば十分 =============================================================== あくまで自然言語ベースである。 主役は人間側の執筆者であり、執筆に注力できるべきである。 エンジニア視点で厳密な構文や配置を決定しない。 固定された構文規則やルールは将来に渡って延々とコストになる。 AI を軸に柔軟な構文解析、言語学的な分析の自動化は進んでいく。 厳密さを要求した後のツケ -------------------------------------------------------- RFC や man ページなど、troff/groff で 編集・更新コストが高くなったものを忘れてはならない。 当時の一部技術者らにとっては A4 に印刷するために役立つツールだったが、 2026 年現在においては存在すら知らない技術者すら多数いる。 一部のパッケージメンテナらが対応してくれているおかげで未だ通用しているが、 それをあえて選び取る必要はない。 最低限のガイドラインとしてのメタデータ表現方法 ============================================================ メタデータの便宜は推奨するが、厳密な構文はあえて避ける。 簡単な ":Refer:" 程度の提案はするが、構文は定めない。 他の候補として Scope/Part, Depends/Require, Refer/Source, Context/Next/Prev などが挙がる。 もちろん英語にこだわる必要もない。:参照: でもよい。 前後関係など別リソース指示にあたっても自然言語を許容する。URL でもいい。 他方でこれは既存のメタデータ記述法の否定でもない。 冒頭の reStructuredText 様のメタデータなど、 既に存在している記載方法や学習コストも執筆時のコストも低いものは排除する必要がない。 ただしその内容評価と利用前提は、若干の場繋ぎが必要にはなる。 最たるものが :License: や :Author: などであって、 それらはできるだけ既存のツールが機械的に読めるようにするべきだろう。 Markdown が好みなら Markdown の様式でやればよい。 LLM などの AI らはアルゴリズムで読んでいない ------------------------------------------------- 現時点の LLM らの数学的基盤から言えば、 :Depend: だろうが #Depend だろうが実質的な差はない。 古典的な構文解析アルゴリズムに囚われる必要はなく、 むしろ意図的に回避することで将来に渡って低コストで柔軟なメタデータ表現をする。 人間にとって把握できる記載にさえすればよい。 URL ベースの参照はできるならするが必須にしない。 参照元を辿っていけることは著作権などの観点から重要ではある。 故に本文書も冒頭に reStructuredText 様式で各種メタデータを記載しているが、 これらは厳密な要件でも提案でもない。 現時点で適していると考えた記法で、ライセンス関連の重要事項を明記しているに過ぎない。 英語版の AI 君の出力を読んでみたら意図がよりわかるだろう。 当時のセッションでは SNS 的な #ハッシュタグ などでも十分機能する事実があると、 メタデータの扱いに関してすら自由度は高くできる可能性を話していた。 #Refer CoreVectorMathematics (in this file, written above). かの出力自体が厳密な構文規則が最早不要であることの証左になっている。 結語 ============= 長らく述べてきたが、要するにヒトが読んでわかる記載でよい。 厳密な構文規則などを決めたくなるのは、 機械的な構文解析手法に慣れ親しんでいる技術者側の発想だ。 これまではそうあるべきだったし、これから若干の期間も必要であるだろう。 ただしそれは移行期であるが故で必要なわけではない。 如何なる構文規則も配置ルールも、 長期に渡って幅広いヒトの執筆者の側に無用なコストを招く。 そもそも Lost in the Middle といった AI らの現時点の問題が、 HTML や CSS やらでノイズまみれで肥大化したデータに一端を発している。 新たなルールをもって制しようという試み自体が、 今後長期的に渡る人類全体へのコストという呪いをもたらしてしまう。 故に llms-full.txt の概念自体に、筆者は全面的に反対する。 情報発信におけるルールも規則も実質なくなれば、全員が利を得る。 執筆側は執筆に注力すればよいだけである。 フォーマットが低コストで柔軟であることが将来に渡っての情報発信を刺激し、 AI らを運用する側はより単純で密度の高い情報源を得られる。 個別サイトで異なる HTML 構造を掘る計算コストとハルシネーションが排除できうる。 執筆してアップロードするという手間をヒトの側において緩和することで、 引用元や参考として列挙できる URL は経時的に増えるだろう。 既存のウェブサイトの在り方は参照元として機能する点や発信の明示性から重要だが、 極端な話 AI との対話セッションの一部自体が新たな情報源になりうる。 ただしその点はプライバシーや権利関係で込み入っているので本文では触れていない。 AI 経由で読者となる側も、つまり大半の人類が、より正しく新たな情報に接していける。 質疑や対話により適切な形式、読みやすい粒度で AI 側が調整して出力してくれる。 古典的な別サイトへの URL リンクとして参照元を検証する必要すらなくなりえる。 大半はチャットセッションに並行して表示されるリソース元を、 AI が適切な粒度に再構成したものを確認すればよくなるだろう。 ネット広告業すらより低コストかつ高効率になりうる。 文脈に沿った表示を行うことで個人情報保護や トラッキングによるプライバシー問題を回避して実装できる。 必要なものはほとんどない。むしろ不要なものを廃するのだ。 ヒトがヒトに分かる情報源を、文書を記せばよい。単純テキストでよい。 旧来の HTML + CSS 経由での発信が好みなら、続ければいい。 これからは余計なものを付け加えないことに価値が見いだせる。 これまでが過剰包装だったのだ。 文脈を理解して学習する AI という存在が現れたことで、 我々はそのコストから開放されうる絶好の機会を得ている。 ならば自然言語で記そう。余計なコストを招く新たなルールは不要だ。