Commit graph

9 commits

Author SHA1 Message Date
Local Dev
8ed051c2da feat(pdf-editor): edit the document's own text on the page, and a line means the whole line
Three complaints, one cause between the first two.

A line of a PDF is rarely one run. pdf.js splits it wherever the file does
— a font change, a kerning adjustment, a colour change — so a heading can
be three spans and an invoice line ten. Replacing the span under the
cursor covered a fragment and left the rest of the line standing, which is
exactly what a replacement that looks like a copy laid over the original
is. A run is now the whole visual line: the spans that share its baseline
and sit close enough to be spacing rather than a second column, with the
spaces the geometry implies put back between them.

And that line is edited on the page. The dialog that used to hold a copy
of the words is gone: the cover goes down first, carrying the line's own
words at the page's own size and colour, and the caret opens on it. The
cover keeps its words hidden while you type, so the original never shows
through the thing covering it. Escape with nothing changed lifts the cover
again and leaves the page as it was found — no mark, no undo step.

The rotate grip was a square like the resize handles, wearing the open
hand that means drag-the-page. It is a disc with a turning arrow now, and
a cursor drawn to match, since no standard cursor means turn. The inline
style that was defeating the stylesheet is gone with it.

Ctrl and the wheel zoom, about the pointer rather than the top-left, so
the words you were reading stay where they were. A plain wheel still
scrolls.
2026-09-27 19:24:28 +02:00
Local Dev
cdad2907cb feat(pdf-editor): the five things you open a PDF to do, named and in front
Select, Text, Highlight, Draw, Sign — labelled, in that order, at the head
of the toolbar. Each one leads a kit rather than hiding it: the shapes sit
behind Draw, underline and strike-through behind Highlight, and nothing is
offered twice.

Draw arms the pen, which is what Draw means when there is one button for
it. The names drop out below 1180px, where five of them would start
pushing the zoom and colour controls off the end.
2026-09-27 18:21:55 +02:00
Local Dev
ce066bb4ca feat(pdf-editor): one select, one text tool, and a signature you can ink and turn
Select does what selecting does in every editor people already know.
Double-click bare page and a caret opens there; double-click the document's
own words and they open for replacement; what is selected copies with
Ctrl+C, pastes with Ctrl+V and goes with Del. Nothing was taken away — a
drag on empty page still gathers an area, and a drag that starts on words
still selects words to copy.

Type text and Edit text were two buttons for one question the click already
answers. They are one Text tool: land on the document's own words and it
offers to replace them, land anywhere else and it starts new text. The
words light up under the cursor so which is which is visible before
clicking, not after.

The toolbar says what it is for. Select, Text and Sign are labelled and set
apart; the drawing kit and the markup kit are their own groups. Sign gets a
pen icon over a signature rather than a squiggle that could have been
anything.

Signatures take ink — black, blue, red, green — chosen while drawing and
kept with the signature, because people sign in a particular colour and it
belongs to the signature, not to whichever swatch was armed. And they turn:
a grip above the box, free rotation, Shift to snap to 15°, for the signing
line that is not square to the page.

Two faults the tests found, both invisible by eye:

The rotate grip was drawn in the right place and could not be grabbed —
the selection bar floats directly above a mark, which is exactly where the
grip sits, and it swallowed every click. The bar now stands clear of it.

Undo would not undo a first rotation. Restoring a mark with Object.assign
copies the keys the original HAD, so a property the drag introduced
survived the restore; the journal then recorded the rotated state as the
state to go back to. Restoring now forgets keys the original never had,
which fixes every future property with the same shape.

Also: building a document from pictures or joins refuses to start a second
one on top of the first, and says so rather than failing quietly.
2026-09-27 15:09:17 +02:00
Local Dev
10c83a1431 feat(pdf-editor): type on the page, keep more than one signature, move what you drew
Three things the editor made you work around.

Text was typed into a dialog and then placed, so you chose a size and a
weight for words you could not see against the page they were going on.
The click now opens a caret where you clicked, in the font, size and
colour the words will have, with the style bar over it; double-clicking a
stamp reopens it in place. Lining a CSS line box up with a PDF baseline is
measured from the font's own metrics, not guessed.

A signature lived in a single slot. There was nowhere to keep initials as
well as a name, nowhere to change the one you had, and reaching for the
tool again simply stamped the first one — which is the same fault three
times: one slot. It is a library now, with redraw, rename and delete, and
the choice is made when the tool is picked up, so placing stays one click.
An existing single signature is carried into it rather than dropped.

The mark you had just drawn could be resized by its handles and not moved
by its middle, because only the select tool let marks be hit-tested at
all. The SELECTED mark now takes a press whatever tool is armed. A press
anywhere else still draws, and an unfilled shape is still grabbed by its
outline — the same rule select has always followed.

Also: words default to dark ink rather than highlighter yellow, which was
unreadable on white and is now impossible to miss, since you watch
yourself type it.
2026-09-27 13:16:07 +02:00
Local Dev
acc433805e feat(pdf-editor): the dock opens the editor, and the editor offers more than one way in
Clicking the dock raised a native file browser, which was the right answer
while the editor had exactly one thing to offer an empty tab. It is the
wrong answer now: a file browser can only ask which PDF, and the answer is
sometimes none of them.

So the dock opens the editor, and the empty editor says what it can do.
The drop zone stays, and learns to read what it is given — pictures become
pages, several PDFs become one document. Beside it sit the three ways in
as buttons.

Not included: compress, which cannot be done honestly without re-encoding
the images, and split, which is the page rail plus Save a copy.

A document built from pictures or joins has never been on disk, so it is
marked unsaved from the moment it opens — otherwise closing the tab would
bin it without asking. An empty editor also stops claiming to hold a file
called document.pdf.
2026-09-23 00:51:45 +02:00
Local Dev
c4f7251860 feat(pdf-editor): convert a PDF to Word, and say what that costs
A PDF does not contain paragraphs. It contains glyphs with coordinates, and
there is no heading, no list, no table and no guaranteed reading order —
only runs of characters that happen to sit near each other. Converting to
Word means working out where the paragraphs were, from geometry. That
inference is the whole feature, and it is sometimes wrong, so this is called
a conversion and never an edit, and the dialog reports what it found before
anything is written.

Lines are grouped by baseline, runs joined with the spaces a PDF only implies
by leaving a gap, and paragraphs ended where the next line sits unusually far
below, is indented, or where the previous one stopped short of the measure.
Headings come from size relative to the body — which is the most common size
on the page, not the average, because a page of 11 pt under a 28 pt title
averages to something that is neither. Bold and italic come from the font's
name, the only place a PDF records them.

What it refuses to fake is as important. A page set in columns is reported,
not silently interleaved. A page with no text says so, and says why: it is an
image of writing, and reading that needs character recognition this editor
does not have. Tables become plain paragraphs rather than an invented grid,
because a wrong table is harder to repair than no table.

The .docx is written here rather than by a vendored builder: a Word file is a
zip of five XML parts, and the subset that can honestly be produced —
paragraphs of styled runs — is about two hundred lines. Vendoring a document
library would have added another megabyte on top of the four pdf.js and
pdf-lib already weigh, to generate markup we would still have to get right.
Entries are stored rather than deflated, which keeps a compressor out of the
add-on; the CRCs are the part that cannot be skipped, since Word calls the
file corrupt rather than naming the part that upset it.

Text replaced in place converts as replaced. Converting would otherwise hand
back the words the user had just edited away.

Checked by taking the output apart — every CRC verified, both XML parts run
through a real parser — and then, because that is still marking my own
homework, by opening the result in the Word editor extension, where mammoth
reads it with none of my code involved.
2026-09-22 22:06:54 +02:00
Local Dev
d625bd25a1 feat(pdf-editor): replace the document's own text, on its own baseline
Until now "editing" a PDF here meant laying things over it. You could put a
word on top of a word, but the document underneath never changed, and the
result read like a sticker because it was one. This adds the thing the word
Edit actually promises: click a line of the document's text, type different
words, and they land where the old ones were, in the old size and the old
colour.

The position and size come from pdf.js's text layer, which has already placed
a span over every run and carries that run's size in unscaled PDF points — so
the size is right whatever the zoom, which reading it off the rendered box
would not be. The colours come from the rendered page, because nothing in the
text API reports them: the background is the average of the most common colour
bucket in the run's box, since type is a minority of the pixels even when it
is dense, and the ink is whatever sits furthest from that background. On the
test fixture it recovers the marker's red exactly.

Two things that look like details and are not. The bucket only chooses WHICH
pixels are background; the colour itself is their average, because rebuilding
it from the bucket index rounds white down to #f8f8f8 and a not-quite-white
patch on a white page is a visible seam. And the cover reaches below the
baseline by a quarter of the font size, because pdf.js sizes its spans to the
em box: cut the cover to the span and every descender in the original line
survives as a little hook under the replacement.

A replacement is a cover plus text, so it is a mark like any other — movable,
resizable, undoable, and rendered on screen from the same numbers the writer
uses, which is what makes the preview trustworthy.

Said plainly in the dialog and again in the save summary: this hides the
original, it does not remove it. The old glyphs are still in the content
stream underneath. Redact is the tool that takes text away, and it says so
too.
2026-09-22 21:49:37 +02:00
Local Dev
0684906f13 feat(pdf-editor): a mark you placed is something you can still work on
Everything the editor put on a page was final. A text stamp could not be
corrected without deleting it and typing it again, nothing could be resized,
and the only way to remove a mark was a Delete key nobody had been told
about — the selection drew a dashed box and offered no action at all. Placing
a stamp also left its tool armed, so the next click stamped a second copy.

Marks are now editable objects. Selecting one gives it grab handles and a
small bar pinned above it: delete and duplicate for anything, and for text an
edit button, a size stepper and bold and italic. Double-clicking text reopens
it for rewriting in place rather than adding a second one. Placing a text
stamp or a signature drops straight back to the select tool with the new mark
live, which is both what people expect and what puts it immediately within
reach of a nudge.

Resizing is one function over every mark type rather than a special case per
kind: a handle drag produces a new bounding box, and the mark is mapped from
its old box into that one. Text scales by font size instead of stretching its
glyphs, signatures keep their aspect on a corner, and lines offer their two
endpoints instead of a box that would let you stretch them in ways you never
aimed at. A whole gesture lands on the undo stack as one step.

Selecting a thin mark used to mean clicking its outline exactly — about one
screen pixel. Each stroked mark now carries an invisible fat copy of itself
purely to catch the pointer.

New marks to go with it: underline and strike-through, which share the
highlight's text-selection geometry and differ only in where the rule sits; a
plain line; and a fill toggle for rectangles and ellipses. Bold and italic
mean three more Helvetica variants embedded at save time, since a PDF treats
them as separate fonts rather than as a style.

Double-click is detected from the pointer stream rather than from a dblclick
listener, because selecting a mark calls preventDefault() on the pointerdown
and that suppresses the compatibility mouse events the browser would have
synthesised the dblclick from.
2026-09-21 03:13:06 +02:00
Local Dev
6b2c4b25c0 feat(pdf-editor): read, mark up and reshape a PDF without leaving the browser
A PDF that needs a signature, a highlight or a page removed currently sends
the user out to a desktop application or, worse, to a web service that wants
the document uploaded first. Both are poor answers for a browser whose point
is that nothing has to leave the machine. This is a full-tab editor that opens
a PDF, marks it up, fills its forms and saves a new copy, entirely locally.

Two engines, vendored rather than installed, because an add-on ships as a
self-contained folder over the signed update channel and nothing runs a
package manager on the way: pdf.js reads and renders, pdf-lib writes. They
share no state. Everything in between lives in PDF user space — points,
origin bottom-left — which is the one coordinate vocabulary both speak, so a
mark survives zooming, rotating and reordering with no conversion table and
save-time needs to know nothing about how a page happened to be displayed.

The page strip is built from pdf.js's PDFPageView components rather than its
PDFViewer, which renders pages in the file's own order and cannot hide,
reorder or individually rotate one — three of the features here. Text layers
are ours and stay attached for every page, drawn or not, because Theseus's
find bar is Chromium's findInPage over the live DOM and a torn-down text layer
is a page Ctrl+F cannot see. Canvases are virtualised; a letter page at 100%
is 3.4 MB of bitmap.

Redaction is the part worth being careful about. A black box over text hides
nothing — the text stays in the content stream and comes straight out of a
copy-paste — so the editor says so in a modal before the tool can be used,
and on save rebuilds each redacted page as an image, which genuinely removes
it. Pages that were not redacted are untouched. Form widgets and links are
kept, since they were never the leak.

Saving never writes over the original: every save reloads the source bytes and
replays the session onto a fresh copy, so a botched save cannot poison the
next one.

Out of scope for this first version: editing the text that is already in the
document, and writing XFA forms back (pdf-lib cannot, so those are fill-and-
print only, and the editor says so on open).
2026-09-20 20:58:21 +02:00