Analysis

What a browser AI lab can teach before you call a model

Explore tokenisation, temperature, top-k and top-p with worked examples, and see what browser experiments cannot prove about a real AI system.

The short answer

A small browser experiment can expose token boundaries, probability transformations and reproducible checks without calling a language model. That helps a team specify a later pilot, but invented scores and toy merge rules cannot establish real model quality, access control, privacy or deployment performance.

Prepared with AI assistance and reviewed before publication. The scores, text and choices below are invented teaching inputs. Their calculations were checked with a deterministic reference, not measured from a language model or a workplace deployment.

Before commissioning a private or on-premises AI pilot, a team should be able to ask precise questions about the proposed system. What turns text into tokens? What changes when temperature changes? Which choices remain eligible after filtering? A small browser experiment can make those mechanisms inspectable without first buying inference access or supplying business documents.

The strategic value is clearer specifications and better questions. A working probability display does not show whether an assistant answers correctly, respects permissions or helps its users. These examples separate mechanism evidence from the additional evidence a real pilot needs.

Set up an experiment you can explain

Use an experiment card with five fields: input, rule, prediction, observed result and unresolved question. Change one setting at a time. Preserve the original input and the order of operations. A screenshot of a pleasing chart is less useful than a table another person can reproduce.

The worked examples here are complete enough for paper or a calculator; a compatible browser lab is optional. The four-score vector is an independent article exercise, not a promised selectable dataset on the Trust Agent labs page. Check the available version before following interface instructions. This article does not assert that newer lab features are live or that their browser isolation has been proven.

Turn four scores into a distribution

Imagine four possible choices labelled amber, birch, cedar and dawn. The labels are conveniences, not token IDs from a particular model. Assign scores [ln(4), ln(3), ln(2), ln(1)], approximately [1.386294, 1.098612, 0.693147, 0]. A score is not yet a percentage. The resulting probabilities describe selection chances, not confidence that an answer is true.

For positive temperature T, calculate weights exp(score / T), then divide each weight by their sum. Here ln is the natural logarithm and exp is its inverse. You can use the supplied weights without first learning logarithms. This is a temperature-scaled softmax; PyTorch's Softmax reference gives the underlying exponential-and-normalisation definition. At T=1, our weights are exactly [4,3,2,1], whose sum is 10. The probabilities are therefore 40%, 30%, 20% and 10%.

Now specify the entire toy pipeline: temperature first; keep the three highest-probability choices and renormalise; then retain the smallest descending prefix reaching at least 75% and renormalise again. These last two operations are our top-k and top-p stages. Hugging Face's generation documentation describes those controls; the order and tie conventions below are explicit features of this example, not a universal provider contract.

Invented scores with T=1, k=3 and p=0.75; percentages rounded for display
ChoiceAfter temperatureAfter top-kAfter top-p
amber40%44.444444%57.142857%
birch30%33.333333%42.857143%
cedar20%22.222222%0%
dawn10%0%0%

Removing dawn leaves weight 9, so amber becomes 4/9. Amber alone is below 0.75; amber plus birch is 7/9, above it. Keep birch, the choice that crosses the threshold. Their final weights are 4/7 and 3/7. Top-p does not mean “remove every choice whose individual probability is below 75%”.

Predict before changing: apply p=0.75 directly to the original distribution instead. The first two probabilities sum to 0.70, so cedar is also needed. Reversing or omitting a stage changes the retained set. Record the order alongside the settings; the numbers alone do not fully specify the experiment.

Explain a change, then look for failure

Keep k=3 and p=0.75, but set T=0.5. The initial weights become [16,9,4,1]. After top-k, the first two cover 25/29, so the final distribution is [64%,36%,0%,0%]. At T=2, the same pipeline instead retains three choices, with final probabilities approximately [38.863141%,33.656468%,27.480391%,0%]. Explain why a temperature change can affect which candidates survive top-p, not just the height of the original bars.

Our T=0 case is a separate greedy rule: choose the highest score, giving amber 100%. Do not divide by zero in the formula. For tied scores, this reference chooses the first listed choice. That convention makes this example reproducible; it does not promise identical tie handling or end-to-end determinism from a hosted service.

A distribution is also different from a draw. With the final 4/7, 3/7 distribution, place amber first on a cumulative interval from zero. A chosen uniform value of 0.50 selects amber; 0.80 selects birch. The less probable choice remains possible. Those values are illustrative inputs, not measured random samples.

Failure checks: inspect whether probabilities sum to one before display rounding, removed choices stay at zero, and the threshold-crossing choice survives. If a run returns a filtered-out choice, inspect its sampling stage. If only rounded percentages disagree at a cutoff, inspect full precision before declaring an algorithm error. A small batch of draws need not reproduce the percentages exactly.

Follow token merging without inventing a model

Now use the exact text ab ab ab, with two ordinary spaces. Start from eight character symbols: [a,b,␠,a,b,␠,a,b]. Here ␠ displays a space; it is not an extra character in the input. Count adjacent symbol pairs, choose the most frequent repeated pair, and merge its non-overlapping occurrences from left to right. Recount after each step. A frequency tie goes to the pair encountered first.

Two learned merges on one tiny training string
StageChosen pair and countResulting symbolsCount
StartNo merge[a,b,␠,a,b,␠,a,b]8
1a + b, three occurrences[ab,␠,ab,␠,ab]5
2ab + ␠, two occurrences[ab␠,ab␠,ab]3

At step two, ␠ + ab also occurs twice. The first-encountered rule selects ab + ␠. Now no adjacent pair repeats, so this toy stops. Notice that a space became part of a symbol. Our example allows whole-string merges; it does not first impose word boundaries.

Freeze the learned rules in order: first a+b → ab, then ab+␠ → ab␠. Apply them to new text ab cab. The result is [ab␠,c,ab], which joins back to exactly the input. This toy preserves the unseen c as its own symbol. Applying existing rules is different from learning new merges on the new text.

Prediction and counterexample: before running on aaaa, count three adjacent a+a pairs. A left-to-right merge makes only two replacements, producing [aa,aa]. Overlapping counts and non-overlapping replacements are different quantities. If a demonstration reports three replacements, reconstruct the positions and challenge the result.

Hugging Face's BPE explanation distinguishes learning merges from applying them and discusses implementation-dependent ties. Our tiny, character-based example illustrates that idea; it is not a faithful tokenizer for every model. Normalisation, word boundaries, byte handling, special tokens and vocabulary rules matter. Do not use its token count to estimate an actual provider's bill or context limit.

Use the result to define a better pilot

These are two separate mechanisms. The merge exercise learns rules and produces symbols; it does not assign them the four invented sampling scores. The sampling exercise transforms supplied scores; it does not learn them. Neither runs a trained neural network producing context-dependent outputs, optimises a language-model training objective or updates neural-model parameters. Calling the exercise “AI” should not conceal those missing parts.

For a private AI proposal, turn each observation into a question. Which tokenizer and revision will the application use? Which sampling settings and processing order are actually supported? What identifies the model, prompt and source documents used in a comparison? How will missing information, denied access and stopping conditions be checked? The toy gives vocabulary for asking; the proposed system must supply its own evidence.

Keep three decisions separate: whether the arithmetic matches its specification, whether the selected model performs the intended task, and whether the complete application can be operated responsibly. An on-premises location does not answer the latter two. Neither does a browser chart establish isolation, confidentiality, permission enforcement or production latency. Use synthetic inputs while these questions remain unresolved.

Your exit artifact: retain the score table, the changed-order prediction, both merge rules, the new-input trace and one failure explanation. Ask another person to reconstruct a result from that record. If they need an unstated convention, add it. Success means the mechanism is explainable and reproducible, not that a business pilot has been approved.

For further practical study, Trust Agent's Level 1 AI introduction supplies background. Its Level 3 tokenizer workbook is a later coding step requiring enough Python to run scripts and read loops/functions. Confirm current availability and prerequisites. The official references linked here were checked on 27 September 2026; software behaviour and documentation can change.

Contribution and ownership: This AI-assisted piece is credited to Mickarle Wagstaff-Irons - Micky Irons, full name Mickarle Sean Junior Wagstaff-Irons. Unified and Trust Agent belong to the Mickai publication family. Mickai's AI readiness programme is an optional commercial route for discussing a proposed pilot, not an independent endorsement, a requirement for learning or a guarantee of security. Confirm its current scope and availability separately.

Questions readers ask

Is a sampling demonstration running a language model?
Not when it only transforms supplied scores. This article's four scores are invented; no trained model produced them. A real model call needs a separate source of context-dependent scores.
Will every tokenizer produce the token trace shown here?
No. The example uses explicit symbol, pair-count, tie and merge rules. Real tokenizers can use different normalisation, boundaries, vocabularies and algorithms.
Does local browser execution prove a private AI system is safe?
No. Correct arithmetic demonstrates that narrow mechanism. It does not establish data handling, permissions, runtime isolation, model behaviour or operational security.
MICKAI®

Published by Mickai LTD.

Unified covers the field broadly and treats Mickai as one example within it. About the journal and the team.

Mickarle Wagstaff-Irons - Micky Irons, full name Mickarle Sean Junior Wagstaff-Irons. Article author. Biography and related work.

Keep reading