Skip to content

RAT-532: Bump tika.version from 3.3.2 to 4.0.0 and migrate charset detection - #714

Open
dependabot[bot] wants to merge 2 commits into
masterfrom
dependabot/maven/tika.version-4.0.0
Open

RAT-532: Bump tika.version from 3.3.2 to 4.0.0 and migrate charset detection#714
dependabot[bot] wants to merge 2 commits into
masterfrom
dependabot/maven/tika.version-4.0.0

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Aug 24, 2026

Copy link
Copy Markdown
Contributor

Bumps tika.version from 3.3.2 to 4.0.0.
Updates org.apache.tika:tika-core from 3.3.2 to 4.0.0

Changelog

Sourced from org.apache.tika:tika-core's changelog.

Release 4.1.0 - unreleased

  • The Kafka pipes iterator no longer stops at the first empty poll. A newly subscribed consumer spends its first poll(s) joining the group and returns empty even when the topic has a backlog, so the iterator could enqueue zero files and report success. It now waits for a partition assignment (bounded by the new assignmentTimeoutMs, default 30s) and requires a continuous quiet window (drainIdleMs, default 1s) before concluding the topic is drained. groupInitialRebalanceDelayMs is deprecated and no longer sent to the consumer: it is a broker setting that Kafka has always ignored (TIKA-4833).

  • Pipes IPC: carry inline document bytes as a raw binary field beside the tuple in the request envelope -- never inside the tuple or its ParseContext -- and disable Smile's 7-bit binary encoding. Tuple JSON serialized by 4.0.0 with an "inline-bytes" parse-context entry no longer loads; it is rejected with a tailored message (TIKA-4829).

  • Digesting embedded documents no longer buffers each embedded object to a temp file. Zip entries are re-read from the parent archive on rewind, and a new process-wide CacheMemoryBudget (seeded by the pipes forked server; default 256MB, clamped to a quarter of the fork's heap; tunable via -Dtika.pipes.cacheMemoryBudgetBytes in the config's forkedJvmArgs, <=0 disables) lets embedded objects stay in memory past the per-object 1MB threshold. New public API on TikaInputStream: get(IOSupplier,...), enableRewind(CacheMemoryBudget), getSeekableByteChannel(). Zip/7z/epub/odf parsing and zip container detection now read through seekable channels, so after detection/parsing a TikaInputStream may no longer be file-backed (hasFile() false); getPath()/getFile() still work and spool on demand (TIKA-4828).

  • Pipes now carries the caller-supplied Content-Type across the worker's fresh-metadata boundary as a soft detection hint, so every forked-parse endpoint (/tika, /meta, /rmeta, /unpack, /async, /pipes, plus tika-grpc and embedded PipesForkParser) can route on a client Content-Type, not only on the filename. Detection keeps the hint only when it equals or specializes the content-detected type (e.g. refining image/tiff to image/x-canon-cr2); for bytes with no magic it can select any type, matching the routing power the filename already had. The CONTENT_TYPE_USER_OVERRIDE key is deliberately not carried, so the hint cannot force an unrelated type (TIKA-4825).

  • OneNote extraction now follows document order, omits superseded page revisions, sorts author metadata, extracts embedded object BLOBs, and bounds malformed-input recursion and file-derived allocations. Parse warnings and embedded relationship IDs are exposed in metadata. Malformed or truncated files that cannot be fully parsed, and files whose walk yields no content, now fall back to the legacy string dump instead of failing or returning empty output. The legacy MS-ONESTORE walker bounds its recursion (depth caps plus file-node-list and fragment-chain cycle guards) and now honors shouldParseEmbedded for embedded file data

... (truncated)

Commits
  • 514e1b3 [maven-release-plugin] prepare release 4.0.0-rc1
  • 4e39e07 revert second rc1 attempt
  • 7975986 javadocs take 42
  • 95d4235 [maven-release-plugin] prepare for next development iteration
  • 666289b [maven-release-plugin] prepare release 4.0.0-rc1
  • c9f6585 TIKA-4808 - revert aborted 4.0.0-rc1 release commits; fix per-module javadoc ...
  • 41183e7 [maven-release-plugin] prepare for next development iteration
  • 5dd7fc7 [maven-release-plugin] prepare release 4.0.0-rc1
  • 4b231cf TIKA-4808 -- prep CHANGES.txt for release
  • 532a685 TIKA-4808 - remove access to the network parser from cli (#3036)
  • Additional commits viewable in compare view

Updates org.apache.tika:tika-parser-text-module from 3.3.2 to 4.0.0

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
  • @dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

Bumps `tika.version` from 3.3.2 to 4.0.0.

Updates `org.apache.tika:tika-core` from 3.3.2 to 4.0.0
- [Changelog](https://github.com/apache/tika/blob/main/CHANGES.txt)
- [Commits](apache/tika@3.3.2...4.0.0)

Updates `org.apache.tika:tika-parser-text-module` from 3.3.2 to 4.0.0

---
updated-dependencies:
- dependency-name: org.apache.tika:tika-core
  dependency-version: 4.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
- dependency-name: org.apache.tika:tika-parser-text-module
  dependency-version: 4.0.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot dependabot Bot added dependencies Pull requests that update a dependency file java Pull requests that update Java code labels Aug 24, 2026
@ottlinger ottlinger changed the title Bump tika.version from 3.3.2 to 4.0.0 RAT-532: Bump tika.version from 3.3.2 to 4.0.0 Aug 24, 2026
* Upgrade to Tika v4.0.0
* Migrate to new charset detection logic of Tika 4.
* It returns the best possible guess as encoding.
* In contrast to v3.x windows-1252 is detected instead of UTF-8/ISO-8859-1,
thus tests had to be changed as well for all UIs.
* Keep the performance optimisation (read only 256 bytes) introduced via RAT-494
@sonarqubecloud

Copy link
Copy Markdown

Quality Gate Failed Quality Gate failed

Failed conditions
66.7% Coverage on New Code (required ≥ 80%)

See analysis details on SonarQube Cloud

@ottlinger
ottlinger requested a review from Claudenw August 24, 2026 21:54
@ottlinger

Copy link
Copy Markdown
Contributor

@Claudenw would you mind starting a review (I'll add the tests to pass the quality build later). The new Tika4 API detects charsets differently, thus so many changes in test expectations. WDYT?

@ottlinger ottlinger changed the title RAT-532: Bump tika.version from 3.3.2 to 4.0.0 RAT-532: Bump tika.version from 3.3.2 to 4.0.0 and migrate charset detection Aug 24, 2026
Comment thread src/changes/changes.xml
-->
<release version="1.0.0-SNAPSHOT" date="xxxx-yy-zz" description="Current SNAPSHOT - release to be done">
<action issue="RAT-532" type="add" dev="pottlinger">
Update to Tika 4.0.0: new charset detection logic in Tika returns different values compared to 3.x before, such as windows-1252 instead of ISO-8859-1.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this is an error. We have some windows-1252 files but most are ISO-8859-1

I think for our purpose we can label windows-1252 as ISO-8859-1. I need to check the list that is returned from the new Tika and see if it includes ISO-8859-1 as one of the encodings. I think we should select ISO over windows when we have the option. This PR needs investigation and work.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Tika4 logics is to return the "best" charset. In contrast to version 3.x this changed into windows-1252. The new implementation returns the first hit. Personally I wouldn't want to introduce new logics on the RAT-side to generalise into ISO-8859-1 and would take the change as tika-induced and document it in our changelog.

Comment on lines +194 to +197
if (results.isEmpty()) {
DefaultLog.getInstance().debug(String.format("No encoding found for file '%s'", documentName));
return null;
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This code does not do the same thing. the debug should be a warning.

And what happend to unsupported character sets?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We do not have an explicit test for unsuppoorted character sets. Tika handles this internally and returns no charset. If no charset is returned RAT will mark as UNKNOWN if I'm not too mistaken.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file java Pull requests that update Java code

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants