Project tools
Analysis report
See the source-word total, internal repetitions, translation-memory coverage, and a transparent workload estimate before you begin a job.
The Analysis report counts the visible words in each source segment, finds repeated source segments inside the document, and checks the active translation memory for the same language pair. It then groups the source words by match quality. The report does not translate anything and does not consume your monthly word allowance. It is part of the Translation Memory workflow (Translator Pro and up).
The match bands
| Band | Meaning | Weight |
|---|---|---|
| Repetitions | The same segment appears more than once in this document. Translate it once, reuse it. | 15% |
| 100% | An exact match exists in your TM. | 20% |
| 95-99% | A near-exact match; usually a small edit. | 30% |
| 85-94% | A solid fuzzy match; partial rework. | 60% |
| 75-84% | A weak fuzzy match; often faster to rework than retranslate. | 80% |
| New | Nothing similar in memory; full translation. | 100% |
The estimated workload is not a second literal word count. It is the sum of each row's source words multiplied by the displayed weight. For example, 4,372 new words at 100% plus a five-word repetition at 15% gives 4,372.75 workload words, displayed as 4,372.8. The 4.2-word reduction is the estimated saving from translating that repeated sentence once and reusing it.
Weights are planning assumptions, not a universal price list. CAT tools use the same weighted-count method, but agencies and tools choose different percentages. DeepReference displays its weights in the report so you can see exactly how the estimate was produced. Your actual quote, editing time, and billing policy remain your decision.
What counts as a source word
Formatting tags and internal layout markers are excluded. Words joined by an apostrophe or hyphen are counted as one token. For Chinese and Japanese source text, each Han or kana character is counted as one word-equivalent because those scripts do not normally separate words with spaces. This count can differ slightly from Microsoft Word or another CAT tool because tokenization rules are not identical across products.
Why a report can show zero TM entries
The TM total covers the active memory for the document's source and target language pair. If it is zero, no matching memory is active for that pair, so the 100% and fuzzy rows will be empty. Internal repetitions can still reduce the estimate. Approving segments in an editor or importing a TMX file builds the memory used by future reports.
How to run it
Open your documents list
Every completed DOCX row has an Analysis button.
Read the report
The modal shows source words, estimated workload, the active TM size, and each band's contribution. No translation is triggered and no words are consumed.
Use both totals when planning
The source-word total describes the size of the document. The workload estimate describes how much of it may need fresh work after repetitions and memory matches are taken into account. Keep both figures in a quote so the client can see what was counted and which weights were used.
Keep exploring
Translation memory
Save approved source and target segments, reuse exact matches, and review fuzzy matches when similar text appears.
Document translation
Upload DOCX, PDF, XLSX, or PPTX files, translate their text, and rebuild the result in the original file type.
Plans and billing
See the current plan limits, included workflow tools, word-counting rules, billing options, and capacity top-ups.