|
Research report
From Evidence to Edits: Rule-Guided Language-Model Video Editing and a Small On-Device Student ModelOctober 2026
AbstractA companion report condensed the evidence on what keeps viewers watching short-form video into hook principles, retention principles, genre templates and length bands [1]. Here we ask whether those findings make automated editing better, and whether a small model that runs on a phone can learn to edit as well as a large one. We give the findings to a language-model editor in two ways, as an ordered reasoning brief and as a deterministic post-processor that enforces the mechanical rules on any plan, and evaluate on fourteen openly licensed recordings (52.9 minutes). Compared with the same editor without the evidence, the evidence-guided editor made edits 20.3 seconds shorter on average (p = .010), passed more of six automatic rule checks (p = .014) and brought the first payoff forward from 24.6 to about 15 seconds; a blind, order-balanced language-model judge preferred its edits on 7 recordings and the baseline's on 3, with 4 ties, a direction too small a sample to confirm (p = .344). We then taught the same material, lesson by lesson, to Senotel Vision-1.0, a 35 KB rule-and-weights editor that plans an edit in about 20 milliseconds without a network connection. It passed the rule checks at least as often as the large editor but did not reach it: it chose the same material far less often than the large editor agrees with itself (F1 0.33 against 0.59), the judge preferred the large editor on 11 of 13 recordings (p = .006), and fitting its weights to the large editor's choices did not beat the evidence-derived priors on held-out recordings. We report both results in full and describe what the negative one teaches. Keywords: automated video editing; short-form video; language models; knowledge distillation; on-device models; evaluation 1. IntroductionAutomated editors can now cut a long recording into a short video in seconds. Whether the result holds a viewer depends on decisions that are easy to state and hard to make: which moment opens the video, what context the viewer needs, what to leave out, and where to end. A companion report reviewed the evidence on these decisions and summarised it as principles labelled by the strength of their support [1]. This report tests whether that evidence improves an automated editor in practice. We study two kinds of editor. The first is a large language model that reads the transcript and selected frames of a recording and returns a plan for the edit. The second is Senotel Vision-1.0, a small, deterministic model designed to run on a phone, which we taught the same material as a fourteen-lesson curriculum and then tried to fit to the large editor's decisions. Our contributions are:
2. Related workComputational editing of speech-driven video. Berthouzoz and colleagues linked transcripts to interview footage and visualised good cut points from the words and the picture together [2]. QuickCut aligned narration with annotated footage and chose frame-level cuts by dynamic programming under film-editing guidelines [3]. Leake and colleagues selected a take for each line of a scripted scene with a probabilistic model of editing idioms that the user can switch on and off [4]. Fried and colleagues edited talking-head video by editing its transcript and re-rendering the face [5], and LAVE placed a language-model agent in charge of planning and executing edits from natural-language requests [6]. These systems help an editor make the edit the editor intends. The editors studied here are asked to choose the story itself, which requires an explicit account of what keeps viewers watching; that account is the subject of the companion report [1]. Our editors cut real footage only and never synthesise speech or faces. Measuring retention. Viewing logs show that duration explains much of how much of a video is watched [7], that interest peaks coincide with visual transitions [8], and that viewers abandon slow starts within seconds [9]; short-video recommenders correct watch time for duration [10–12]. None of this measures what an editor can change inside a short video, which is why we rely on a proxy and a judge, and why we treat both with caution. Language models as judges. Using a model to compare two outputs is cheap but biased: judges favour a position, longer answers and their own outputs [13]. We follow the published remedy of judging each pair in both orders and counting only verdicts that agree, and we use a judge from a different developer than the editor. 3. From evidence to an editor3.1 The value-of-continuing proxyThe companion report proposes that a viewer keeps watching while the expected value of the next few seconds exceeds that of the next video in the feed [1]. We implement this as a discrete-time hazard over the edit, in steps of 0.25 seconds. At each step t the hazard of leaving is h(t) = b(t) · exp( −(1.4 Q + 1.0 min(1.2, E) + 0.6 P) + 0.6 C + 0.3 max(0, B − 3) )
where Q is the strength of the open question (set by the hook, refreshed by lines that open a new gap, decaying with a time constant of max(20 s, 0.6 × length) and capped low in the final 15%); E is recent value (each sentence carries a value from 0.05 for a greeting to 1 for a laugh or a high-stakes line; cuts, cutaways and on-screen text add smaller amounts; decaying with a 2.5 s time constant); P is the pull of progress, stronger when the content is visibly numbered; C is processing cost (more than two new elements at once, or speech faster than four words per second); and B is the time since anything new, with silence counting one and a half times. The base rate b(t) is 0.20 per second in the first second, 0.10 until three seconds and 0.035 thereafter, and half as high again if the first word comes after 0.6 seconds, which encodes the stay-or-swipe window [9, 14]. Survival is S(t) = exp(−∫h), and the score reported here is 100 × (½ S(T) + ½ mean S), an average of estimated completion and estimated share watched. The proxy is not a measurementIts constants express directions and rough magnitudes taken from the evidence. They were fixed once, before any comparison, and never tuned, but they have not been fitted to any audience. We use the proxy only to locate weak spots and to compare edits of the same footage, and Section 7 explains why a higher score is weak evidence of a better edit.
3.2 Six automatic checks
Table 1. The automatic checks. The band targets for RQ2 and R1 are those of the companion report, with one second of tolerance on R1. 3.3 The language-model editorThe editor receives the transcript of a recording with word timings, a small set of frames and audio measurements, and plans the edit in two passes: a story pass that decides what the video is about and how it is told, and an edit pass that turns the story into clips, on-screen text, cutaways, sound and music. In our experiments it was a Gemini Flash-family model served through the Gemini API [21]; the service substitutes a smaller model of the same family when it is busy, and the model used for each individual run was not recorded. The baseline brief described a good edit in general terms. The evidence-guided brief changes three things. An ordered reasoning brief. The story pass must reason in a fixed order: the genre; the length band; several hook candidates, each with its type, the gap it opens and where the footage pays it off; one hook chosen by the hook principles; then the beats laid out on the genre template with the main question and the smaller ones. It returns the genre, band, hook type, gap and payoff time, which are validated. The edit pass receives the retention principles and must report where the payoff lands. The brief is generated from the same data as the checks, so the editor and the evaluation quote one source. A deterministic post-processor. Whatever plan comes back, a post-processor applies the mechanical principles: a greeting at the very start and a sign-off at the very end are cut; the first word is brought to within 0.5 s of the start; the edit ends within 0.7 s of the last word unless sound continues; and any silence longer than 1.1 s inside a clip is closed, keeping 0.45 s of breath, unless the gap is filled with sound such as laughter. Runs of jump cuts alternate a held 12% zoom, so that a continuation of the same shot reads as intended and the change acts as a small orienting cue [22, 23]. Shared evidence. Both passes are given the companion report's principles in condensed form, with their evidence grades. 4. Senotel Vision-1.0A cloud editor needs a network connection and a service that may be busy. Senotel Vision-1.0 is designed to run on a phone, offline, within a small download. These constraints rule out a large language model on the device today, so Vision-1.0 is a deterministic program whose every decision is a rule or a weighted judgement traceable to the evidence. 4.1 A curriculum of fourteen lessonsVision-1.0 was taught in fourteen lessons, from what a thought is to planning a whole edit by simulating the viewer. Each lesson has the knowledge it carries, the skill that turns it into a decision, and an automated exam the skill must pass.
Table 2. The curriculum. Passages, not sentences. An early version picked the best sentences one by one and produced edits of fifteen fragments. Viewers absorb a story through a plot they can follow [31] and a coherent order [32], and every jump is a cut across which the viewer must re-orient [30]. The body of the edit is therefore chosen by a dynamic programme over the thoughts in their original order that maximises Σ value(kept) − μ · seconds − κ · jumps, with κ = 0.6 and μ, the price of a second, found by bisection so that the edit fits its length budget. A punchline always brings its setup. 4.2 TrainingAfter the curriculum, the weights of two skills were fitted to the evidence-guided editor's decisions on the recordings of Section 5. The choice of opening thought is a choice of one among many, so it is fitted as a conditional logit with the editor's opener as the positive; whether a thought is kept is fitted as a logistic model, a thought counting as kept when at least half of it lies in the editor's clips. Both are maximum a posteriori fits with a Gaussian prior centred on the evidence-derived weights rather than on zero, so that thirteen videos can move a weight only as far as they provide evidence for. All accuracy figures are leave-one-video-out. Edit length was calibrated separately, on the development recordings only: there, the large editor's median edit was 1.29 times Vision-1.0's, so each genre's length target was multiplied by 1.29. 4.3 Size and speedVision-1.0 is a single 35 KB script (13 KB compressed) on top of a 33 KB module of rules that it shares with the large editor. Planning one edit took a median of 19 ms (maximum 30 ms) across the fourteen recordings on a desktop computer. 5. Experimental setupFootage. Fourteen recordings from Wikimedia Commons, 52.9 minutes in total, chosen to span the genres of the companion report: two personal stories, two stand-up sets, three science explainers, an interview, three tutorials (one with almost no speech), a day-in-the-life vlog, a talk and a product review. All are public domain or licensed under Creative Commons licences permitting reuse with attribution (Appendix A). Each recording was transcribed once by Whisper tiny (English) [37], run locally through Transformers.js [38], and analysed for loudness and motion; all systems received exactly the same input, transcription errors included. Systems. Baseline: the language-model editor with the baseline brief and the earlier post-processing. Baseline + post-processor: the same baseline plans through the new post-processor, isolating the effect of enforcement with the plan held fixed. Evidence-guided: the evidence-guided brief and post-processor, run twice on every recording (runs 1 and 2), because the model samples and one run is one draw. Senotel Vision-1.0.0: Vision as frozen before the held-out evaluation. Senotel Vision-1.0.1: one defect corrected after the held-out evaluation (Section 6.4). Split. Seven recordings form the development set and six the test set (Appendix A). Vision's design changes and its length calibration used the development recordings only; the weight-fitting experiments used all thirteen recordings with speech, always leave-one-video-out; Vision-1.0.0 was frozen before its test-set edits were judged. The recording with almost no speech is outside Vision's scope, as it declines to plan footage without speech. Measures. Length; the six checks; the proxy, with the longest lull and the time of the first payoff as it estimates them; and agreement with the evidence-guided editor's second run, as the F1 score of the source seconds both edits keep and whether both open on the same moment (first clips starting within 2 s of each other, or overlapping). Statistics. Paired comparisons over the same recordings use an exact sign-flip permutation test on the mean paired difference, with Holm adjustment across the four measures of each comparison. With fourteen recordings these tests detect only large effects. Blind judge. For an independent view of quality, each pair of edits of the same recording was rendered as text exactly as a viewer would meet it (the opening text, every spoken line in order, the cutaways, the music and the length) and given to gpt-oss-120b [39], an open-weight model from a different developer than the editor, served by GroqCloud [40]. The judge was asked which edit a typical viewer scrolling a feed would be more likely to watch to the end and be glad to have watched, told that both editors had the same material and not to prefer an edit for its length, and allowed to call a tie (Appendix D). Every pair was judged in both orders at temperature 0; a win counts only when both orders agree, and a split counts as a tie. A second judge from another developer was planned but could not be run, because its hosted service refused every request with rate-limit errors. 6. Results6.1 What the evidence changed in the large editor
Table 3. Means over the recordings each system planned. Proxy: the value-of-continuing score (0 to 100, uncalibrated). Checks: of the six in Table 1. Longest lull and first payoff are the proxy's estimates. Averaged over its two runs, the evidence-guided editor made edits 20.3 seconds shorter than the baseline (p = .010, Holm-adjusted = .040), passed 0.46 more checks (p = .014, adjusted = .042) and scored 9.5 points higher on the proxy (p = .020, adjusted = .042). The longest lull did not change (p = .964). The two runs differ: run 1 improved on all four measures after adjustment, run 2 only on length and proxy and only before adjustment, and run 2 produced the longest edit in the study (a 152-second tutorial). The first payoff moved earlier, from 24.6 s to 13.8 s and 16.1 s. The post-processor alone, applied to the baseline's unchanged plans, shortened edits by 3.3 seconds (p = .004), raised the proxy by 3.3 (p = .008), brought the first word to within 0.02 s on average, and made every edit pass H1, H1b and R10 (Table 4). Most of the change in length, however, comes from the editor's own choices: the evidence changed what it decided to keep.
Table 4. Recordings passing each check. For every language-model configuration, the two checks failed most often are R1 (a long stretch without anything new) and RQ2 (an early payoff).
Table 5. Paired comparisons. The mean paired difference is the first system minus the second; for length and lull, negative values are shorter. 6.2 What a blind judge preferredThe judge returned a verdict for one side in every one of its 81 individual judgements and gave the same verdict in both orders for 32 of 40 pairs (80%); the remainder count as ties. It chose the edit shown first in 51% of its verdicts, so position did not drive it, and in the 32 pairs it decided the shorter edit won 14 times, so it did not simply reward brevity, although some of its stated reasons mention it.
Table 6. Blind pairwise judgements. A win requires the same verdict in both orders; "split" counts pairs on which the two orders disagreed, scored as ties. Exact two-sided sign test on decided pairs. Evidence-guided against baseline. Comparing the two complete systems on all fourteen recordings, the judge preferred the evidence-guided edit on 7, the baseline on 3, and tied 4 (p = .344). The direction agrees with the checks; the sample does not allow a stronger statement. The judge's reasons for the wins named tighter pacing and the removal of filler. Of the three losses, one was the 152-second tutorial produced by run 2, where the judge preferred the baseline's 114-second cut; one was a personal story, where it preferred the baseline's "clearer, more uplifting narrative"; and one was a science explainer, where it preferred the baseline's question-led opening, a preference that runs against randomized evidence on question headlines [17]. Large editor against Senotel Vision-1.0. The judge preferred the evidence-guided editor on 11 of 13 recordings and Vision-1.0.0 on 1 (p = .006). Against Vision-1.0.1 the count was 9 to 1 with 3 ties (p = .021): the correction in Section 6.4 narrowed the margin without changing the conclusion. Vision's single win was a science explainer. Reading the losing edits shows where Vision went wrong. On two personal stories it withheld music and cutaways under its rule for sensitive footage, where the large editor scored both with piano, and the judge named the music and the cutaways in its reasons; on a guitar lesson it laid music over the instrument being taught; and most often the judge found the large editor's edit clearer and more coherent. 6.3 How far Senotel Vision-1.0 gotVision-1.0 passed the checks at least as often as the large editor (5.00 of 6 against 4.36; all six on 4 of 13 recordings) and scored higher on the proxy (58.5 against 42.9). Neither number shows that it edits well: it selects its edit by maximising that same proxy among six candidates (Lesson 14), and it was taught the same principles the checks test. What it does not do is choose the material the large editor chooses.
Table 7. Agreement with the evidence-guided editor's second run. The first row is the ceiling set by sampling: two runs of the same editor on the same footage. Two runs of the same large editor share 0.59 of their material and open on the same moment in 10 of 14 recordings. Vision-1.0 shares 0.33 and opens on the same moment in 3 of 13; on the held-out recordings, 0 of 6. The baseline agrees with the evidence-guided editor about as much as the latter agrees with itself (0.59), suggesting that language models converge on the material of a story even when they cut it differently; Vision-1.0 does not reach that material. Training did not help. Fitted to the large editor's choices, the hook model ranked the editor's opener first in 1/13 recordings and within its top three in 3/13; the evidence-derived priors alone placed it in the top three in 6/13 (Table 8). Across prior strengths, held-out ranking improved steadily as the fitted weights were pulled back toward the priors and never surpassed them (Table 9). The keep model's held-out discrimination (AUC) moved from 0.52 to 0.58 against one run and from 0.58 to 0.56 against the other: no reliable gain. Vision-1.0 therefore retains the evidence-derived priors for both skills, and only the length calibration.
Table 8. Learnability, leave-one-video-out, with each run of the evidence-guided editor as teacher. "Same opening (share)" is the fraction of the thirteen recordings on which Vision's whole edit opens on the teacher's moment.
Table 9. Strength of the prior (λ, the weight of the pull toward the evidence-derived weights) against held-out accuracy, with run 2 as teacher. 6.4 A defect found in the held-out evaluationReading Vision-1.0.0's held-out edits revealed a defect the metrics did not: its opening text was sometimes a fragment, because it cut a long clause down to its densest eight words ("Usually steam on high for around six") or joined two recognised lines ("Feel that beat To celebrate your birthday"). Version 1.0.1 prefers a whole clause of up to ten words, then the clause without its filler, article and softeners, and only then the clause up to a phrase boundary; it never crosses two recognised lines and never ends on a number separated from its range. The change affects only the opening text, not which material is chosen, so lengths and agreement are unchanged. Because it was made after seeing held-out output, results for 1.0.1 are reported separately and are not held-out evidence. Appendix B lists the opening text of every system on every recording. 7. DiscussionWhat the integration shows. Giving a language-model editor the evidence as an ordered reasoning brief changed what it kept: edits became shorter, the first payoff arrived earlier and more checks passed. A deterministic post-processor that enforces the mechanical principles improves any plan, including one made without the evidence. The blind judge points the same way, seven recordings to three, but fourteen recordings cannot confirm the direction, and the judge is a language model reading text, not an audience. Why the proxy cannot settle quality. The proxy was written from the same findings that the editor was given and that Vision was taught, and Vision selects its edit by maximising it. A system that follows the principles will score well on a model of those principles whether or not viewers prefer its edit. This is why we never use the proxy here as evidence that one system edits better than another, why we ran the blind judge, and why calibrating the proxy against real audience retention is the first item of future work. Why Senotel Vision-1.0 falls short. The gap lies in one decision more than any other: which moment is the story. The large editor reads meaning; Vision reads words, timing, delivery and what a recording keeps returning to. That is often enough to find the subject but rarely enough to find the moment: the large editor's opener was among Vision's three strongest hook candidates in 6/13 recordings but ranked first in only 1/13. In the one recording where it ranked first, Vision's simulation step still chose another hook, because the proxy preferred that edit. Thirteen decisions by a teacher cannot teach semantics to a linear model over hand-made features; Table 9 shows the data pulling the weights away from the evidence rather than toward the teacher. Closing the gap requires far more teacher decisions and a representation of meaning, such as a compact sentence-embedding model, and preferably both. A note on the teacher. An editor that a student is trained to imitate should itself be shown to be good by people, not only by proxies. The evidence-guided editor improved on its baseline, but in informal use its edits still showed weaknesses that our measures do not capture: lines cut where a speaker's thought was not finished, context too thin for a newcomer to follow, and cutaways that illustrated a phrase literally rather than its meaning. A student can be no better than the examples it learns from, so the choice of teacher deserves its own evaluation with human viewers before any distillation at scale. 8. Limitations
9. Conclusion and future workEvidence about attention can be given to an automated editor in a form it follows, and doing so moved its edits in the direction the evidence recommends. A small model taught the same evidence follows the rules but does not yet choose the right moments; that requires a sense of meaning which rules over words do not provide. Our next steps are, in order: to choose the teacher by human evaluation on creators' own footage; to collect a much larger set of the chosen teacher's decisions on openly licensed recordings; to give Senotel Vision a compact representation of meaning that runs on the device; and to calibrate the value-of-continuing proxy against real audience-retention curves shared with consent, so that future comparisons can rest on viewers rather than on models. Ethics, licensing and disclosure. All footage is in the public domain or under Creative Commons licences permitting reuse with attribution; authors and licences are listed in Appendix A. The footage was used for evaluation only and is not redistributed in edited form. No human participants were involved. One recording is a personal account that includes bullying and suicidal thoughts; it is referred to here only by its licence information, and the editors studied are designed to treat such footage with restraint (no comic sounds, no music by default). The research, code, experiments and this report were produced with the assistance of an AI system (Claude, by Anthropic) under the direction of Senotel. The editor in the experiments used Google's Gemini models [21]; the judge used OpenAI's open-weight gpt-oss-120b [39] served by Groq [40]. Senotel develops automated video-editing tools; no external funding was received. References
Appendix A. Footage
Table A1. The fourteen recordings, linked to their Wikimedia Commons pages, with licence and author as credited there. "Label" is the genre assigned when the footage was chosen; the systems detect genre themselves. Appendix B. Opening text by system
Table B1. The text each system placed on screen at the start. Transcription errors are reproduced as the systems saw them. Appendix C. Per-recording results
Table C1. Length (s) / proxy score / checks passed (of 6), per recording and system. "No plan": Vision declines footage without speech (Section 5).
Table C2. The judge's verdict per recording ("split": the two orders disagreed). Appendix D. The judge's instructionsTwo editors each cut one short vertical video (for TikTok, Instagram Reels or YouTube Shorts) from the SAME raw footage. You cannot see the videos, so each is described as a viewer would experience it: the text on screen at the start, every spoken line in the order it is heard, the cutaway pictures, the music and the length. The spoken lines come from automatic transcription, so expect some misheard words in both. Picture a typical viewer who comes across each video while scrolling. Which edit are they more likely to keep watching to the end, and be glad they watched? Judge the editing, not the footage: both editors had exactly the same material. Do not prefer an edit just because it is longer or shorter. If neither is clearly better, answer tie. EDIT A EDIT B Reply with JSON only, nothing else, in this shape: {"winner": "A" or "B" or "tie", "reason": "one sentence"} Each edit was given as: Length; Text on screen at the start; What is said, in order (numbered lines); Cutaway pictures; Music; Sound effects (a count). Reproduced verbatim except for its one-line title, which names an internal service and is omitted. |