Skip to content
FunDev
FunDev
ai

OKF: I Couldn't Find a Reason to Adopt It Yet

OKF: I Couldn't Find a Reason to Adopt It Yet
30 views
14 min read
#ai

In my previous post, I introduced a way to organize an organization's knowledge with OKF. This time, I tried it in an AI report for sellers that had been generated from ordinary Markdown. Across 180 runs, I compared three approaches: providing the original documents, summarizing them by concept, and automatically selecting the concepts needed.

With the current volume of documents and the way we use them, I couldn't find a reason to adopt OKF right away in place of ordinary Markdown. Summarization sometimes reduced the input, but I found no advantage in the reports' judgments or preservation of necessary content substantial enough to justify changing the existing approach. Factoring in the operational burden of preprocessing and retrieval, I decided that keeping ordinary Markdown was the better choice for now.

The time to test it again is when the text becomes much longer than it is now. Once feeding in all the documents at once becomes difficult, OKF may help as an index that points to the knowledge needed. That possibility is the next question to test with more documents, rather than a result established by this experiment.

Approach to use nowOrdinary MarkdownKeep the existing documents and delivery method
Decision on OKF adoptionOn hold for nowNo practical gain found to justify the switch
Condition for another testMore documentsCheck whether it becomes useful as an index

01. Three Inputs, One Pipeline

The report is HTML that reads a seller's product, social media, and marketplace status, then suggests changes in the form of cards. Since a generation pipeline already existed, there was no need to build three copies of it. I swapped only the input at two points: page_sources, which fetches the sources, and page_payload, which builds the request body.

Source Markdown documents
Input method 1Documents as they are

Concatenate the document list specified for each page.

Input method 2Summaries by concept

Summarize for the generation prompt and cache the result.

Input method 3Selected content

Select paths from the concept index in the DB and pass along their content.

Shared request assembly → LLM generation → HTML validation → Storage
The reference knowledge is what changes. The generation engine's contract and downstream processing path are shared.

Summary Bundles Went into the DB, Not a Folder of Files

I sent the source documents and the generation prompt for the relevant area to a preprocessing LLM together. I parsed the returned result into a {path, content} array and stored it in a single DB row. Logically, these are multiple concept files, but the implementation does not create an actual file tree.

In one case, creating the initial summary took about 302 seconds. Shorter input therefore does not immediately mean a lower total cost. How much the cache is reused matters. The cache key included the source hash and preprocessing prompt identifier, but changes to the generation prompt were missing from the invalidation conditions. Since the summary is tailored to the generation objective, this also needs to be managed in production.

This implementation takes inspiration from OKF to create summaries by concept. I instructed it to use frontmatter, but did not attach a conformance validator. The experiment also did not include using trust metadata such as verified and stale_after.

Concept Retrieval Became a Separate Subsystem

The retrieval approach splits the source into concepts and loads them into the DB, then reads an index and selects the paths it needs. Documents that need verbatim preservation are stored without conversion. I also built a path for comparing the selections against reference paths and displaying metrics in an admin interface.

The system I builtConcept retrieval → Report generation
Source documents are converted or preserved verbatim in a concept corpus, then pass through concept selection and shared report generation before the HTML is stored. The dashed path comparing reference paths with selection records was not executed during the experiment.
A public-facing reconstruction of the supplied Archify flowchart. Dashed lines mark the scoring path that was not executed in this experiment. On small screens, scroll horizontally to read the diagram.

The retrieval approach reused the selection cache in all 60 runs. What this experiment repeatedly measured was therefore reports generated from six inputs that had each been selected once, rather than a system performing a fresh search on every run. Because the implementation skips comparison with the reference answers when reusing the cache, retrieval precision and recall were empty in all 60 runs. I measured output repeatability, but could not directly score retrieval at its core job.

02. What the Number 180 Contains

The experiment ran on September 2–3, 2026. I generated each of six requests—products, social media, and marketplaces for two sellers—ten times using each of three input methods. Below, I call a group with the same request and input method a “cell.” There are 18 cells in total.

I fixed the model and reasoning effort and turned off additional refinement and candidate screening. For the original-versus-summary comparison, I alternated the methods within each page. Concept retrieval was added in a separate run, so this was not an experiment in which all three methods were randomly assigned together. In 3 of the 6 retrieval cells, the source documents came from a different point in time than those used by the other methods.

Distinguishing the Plan, Stored Records, and Analysis Sample
CategoryCountMeaning
Planned repeated runs180 runs2 × 3 × 3 × 10
Stored execution records182 recordsIncludes a recovered extra result and a timeout record
Records containing HTML181 recordsOriginal 61 · Summary 60 · Retrieval 60
Used to calculate stability173 recordsExcludes 1 extra result, 6 source changes, and 1 collapsed output

Thus, “180 runs” in this post refers to the size of the experimental design. The stability table uses 173 records; the search for suspicious text uses the 181 records with surviving HTML. One timeout had been omitted by the filename pattern used in the earlier aggregation. That must not be read as evidence that there were no failures.

What the Request Byte Comparison Verified

I reassembled requests using the same rules as the actual pipeline and compared them against the call logs. There are 118 stored verification records. Of these, 97 had untruncated logs with both system and user messages matching, 19 had truncated user logs, and 2 had mismatched user messages. This check helped assess the comparison conditions, but did not fully prove input control across all 180 runs.

03. Similar Words, or the Same Content?

The first metric is a mechanically calculated Jaccard similarity of word sets. After stripping HTML tags, I formed word sets, compared every pair of outputs within each cell, and took the average. A cell with ten outputs remaining has 45 pairs. Those pairs share outputs, so they do not represent 45 independent experiments.

Jaccard similarity = Number of shared words ÷ Total number of unique words × 100

I compared sets of card titles in the same way. Word similarity reflects vocabulary repetition, while title similarity reflects part of the layout. Neither directly measures factual accuracy or content completeness. Similarity can be high even if the exact same item is omitted every time.

The second measure is a content review. I removed tags, classes, and image URLs, leaving a “skeleton” of titles, judgments, evidence, checks, and table rows for an LLM reviewer to read. It evaluated five dimensions—judgments, suggestions, numbers, checks, and structure—and extracted items that stayed the same and items that changed. In a separate verification pass, the resulting 134 discrepancies were classified as 96 confirmed, 37 partially confirmed, and 1 unconfirmed. These results were not exhaustively verified against human-authored reference answers.

Limitations of the Calculation

The tokenizer splits only on whitespace and some punctuation. Commas and periods are also separators, so numbers can be split apart, and word order and frequency are ignored. Titles are extracted using a particular HTML tag pattern, so markup changes can affect the result even when the meaning stays the same. The figures in this post were recalculated from the original HTML using the existing rules.

04. Summaries Led to More Repeated Words

For word similarity, the summary bundle scored higher than the original in all six requests. The cell averages were 62.4 for the original, 72.6 for summaries, and 59.1 for concept retrieval. The average difference between summaries and the original was 10.2 percentage points.

0–100 · Higher means more similar word choicesWord similarity for the same request with three input methods
Original Summary Retrieval
Seller A · Products
Original85.3
Summary92.1
Retrieval82.7
Analysis sample · Original/Summary/Retrieval: 8 / 8 / 10
Seller B · Products
Original85.4
Summary87.1
Retrieval87.7
Analysis sample · Original/Summary/Retrieval: 9 / 9 / 10
Seller A · Social media
Original63.5
Summary73.5
Retrieval50.7
Analysis sample · Original/Summary/Retrieval: 10 / 10 / 10
Seller B · Social media
Original58.1
Summary68.7
Retrieval46.3
Analysis sample · Original/Summary/Retrieval: 10 / 9 / 10
Seller A · Marketplace
Original38.7
Summary57.2
Retrieval43.3
Analysis sample · Original/Summary/Retrieval: 10 / 10 / 10
Seller B · Marketplace
Original43.1
Summary56.8
Retrieval43.9
Analysis sample · Original/Summary/Retrieval: 10 / 10 / 10
The average Jaccard similarity of word sets across output pairs in each cell, calculated from 173 analysis samples. Some retrieval cells use a different generation of source documents, so the ranking of the three methods cannot be interpreted as a causal effect.

These results mix the effects of input length, summarization, and more uniform phrasing. Since this was not an experiment that changed only the format while keeping the content identical, the increase cannot be isolated as an effect of OKF syntax itself.

Differences between request areas were also substantial. The average was 86.7 for products, 60.1 for social media, and 47.2 for marketplaces. The spread between area averages was 39.6 percentage points, compared with 13.5 percentage points between input-method averages. Across these six requests, the observed differences between request areas were larger than those between input methods.

Similar Words Do Not Guarantee the Content

Average judgment-consistency scores were similar: 90.0 for the original, 89.8 for summaries, and 89.2 for retrieval. Score variation during reevaluation was greater than these differences, making it difficult to use these scores to determine which method produced more stable judgments. The effects I could confirm were less input and more similar wording. Those alone were not enough to conclude that the output had improved sufficiently to justify replacing the current Markdown approach with OKF.

05. Evidence and Structure Changed More Often than the Direction of Recommendations

The item lists extracted by the reviewer were more interesting than the scores. I regrouped by content the 312 items classified as “always the same” and the 216 classified as “different across runs.”

Composition of the item lists extracted by the reviewerFixed items ÷ (Fixed items + Variable items)
Judgments and suggestions79.1%
144 fixed · 38 variable · 182 total
Checks and items to retain64.3%
83 fixed · 46 variable · 129 total
Supporting figures43.5%
64 fixed · 83 variable · 147 total
Layout structure30.0%
21 fixed · 49 variable · 70 total
These are neither repeated-run success rates nor accuracy scores. The four categories contain different numbers of items, and the results depend on how the LLM reviewer divides the content into items.

The share classified as fixed was 79.1% for judgments and suggestions, and 43.5% for figures. This does not mean that “the judgment was reproduced eight times out of ten.” These are proportions within item lists with different denominators. The values also change depending on how finely an item is divided.

Even so, this showed where to pay attention when reading the reports. Recommendations could point in a similar direction while the supporting numbers and additional checks changed. Actions were not always fixed either: 4 of the 14 severe discrepancies concerned judgments and suggestions.

06. Other Problems Found During Repeated Runs

Apart from the input-method comparison, defects in generation also stood out. One output had 753 characters and 0 cards, yet was saved as a success. The other nine outputs in the same cell averaged 10,371 characters and all had 6 cards.

Average of the other 9 outputs in the same cell10,371 characters · 6 cards
Collapsed output saved as a success753 characters · 0 cards

Some results also contained strings such as ennials or traces of intermediate work in the middle of sentences. Searching all 181 HTML outputs for infrequent Latin-script tokens produced 23 candidates for another review. All three methods had them: original 5/61, summary 10/60, and retrieval 8/60. Since rare but legitimate words could also be included, I did not treat these figures directly as contamination rates.

These defects need to be addressed separately in the generation pipeline. I changed the way knowledge was delivered in this experiment, but that change alone did not make the output dependable enough to use without concern.

07. Keep Ordinary Markdown for Now

Summarization clearly had a compression effect. In the product area, it reduced input by about 60–65%. But for inputs that were already short, such as social media, it made them larger.

Original → Summary bundle · Comparison of stored input character counts
RequestOriginalSummaryChange
Seller A · Products81,782 characters32,699 characters-60.0%
Seller B · Products87,888 characters30,563 characters-65.2%
Seller A · Social media24,567 characters27,155 characters+10.5%
Seller B · Social media24,237 characters25,483 characters+5.1%

The table shows input character counts. Summarization adds LLM calls and cache management, so fewer characters do not immediately mean lower total cost. Above all, being able to compress the input is not by itself a reason to adopt OKF. The criterion should be whether the benefit is substantial enough to change what is currently done with Markdown.

Concept retrieval did not support making the switch either. Compared against key output items from the other methods, only 18/46 or 25/51 items were preserved, and 12 of the 14 severe discrepancies came from retrieval. These figures are comparisons between output items, rather than recall measured against a retrieval answer set.

The retrieval implementation I built comprised 5,820 lines across production code, tests, and DDL, with 20 routes and 4 tables. That scale reflects the maintenance burden of my implementation, rather than a requirement of the OKF specification itself. Without a clear advantage over the existing approach, there was no reason to add that burden to production.

The conclusion of this experiment is therefore to keep ordinary Markdown in the current service and put OKF adoption on hold. Whether summarization or conceptualization is possible and whether it needs to be adopted now turned out to be separate judgments.

08. Test It Again as an Index When the Documents Get Longer

Currently, I specify and pass along the list of Markdown documents needed for each request. If the number and length of the documents keep growing, managing that list or supplying all the source text at once may become burdensome. At that point, an index that locates the necessary concepts and their source passages could become useful.

What I want to establish is whether an index built with OKF can find all the knowledge needed with less input. As document volume increases in stages, I need to compare passing ordinary Markdown directly, adding a simple table of contents to Markdown, and using an OKF concept index. This would distinguish the effect of having an index from the effect of choosing OKF.

If that comparison shows practical gains in preserving necessary items, input volume, processing time, and cost, adoption can be reconsidered. Because this retrieval experiment repeatedly reused cached selections, it did not establish how well the index could find the right material as the number of documents grows.

Related posts

Comments

Korean and English pages share this conversation.

Write a comment

0 / 5,000
You will need this password to edit or delete this comment.

Loading comments…