Apache Tika
Revision as of 10:59, 7 November 2014 by Prwheatley (talk | contribs)
The printable version is no longer supported and may have rendering errors. Please update your browser bookmarks and please use the default browser print function instead.
Description
Java based tool for identifying file formats using signatures and extracting metadata and text content from documents.
User Experiences
- Comparing how Apache Tika and DROID perform HTML identification: How much of the UK's HTML is valid?
- Apache Tika is a core component of the Web Archive Discovery indexer and profiler.
- A number of pages on the OPF Wiki mention Tika.
Development Activity
Error in widget Ohloh Project: unable to write file /var/www/html/extensions/Widgets/compiled_templates/wrt660678eb4812d5_25200024
Release Feed
Link to any RSS feed that is updated when new releases occur, if any, e.g: Failed to load RSS feed from http://projects.apache.org/feeds/rss/tika.xml: There was a problem during the HTTP request: 404 Not Found
Activity Feed
Link to any RSS feed that is updated when issue or code updates occur, if any, e.g:
- 2024-03-29 08:16:29
- ASF GitHub Bot updated a link from ASF GitHub Bot logged '10m' on Roman Puchkovskiy started progress on Roman Puchkovskiy changed the Priority to 'Blocker' on
- by Roman Puchkovskiyhttps://issues.apache.org/jira/secure/ViewProfile.jspa?name=rpuchrpuchhttp://activitystrea.ms/schema/1.0/person
- 2024-03-29 08:16:15
- Roman Puchkovskiy changed the Assignee to 'Roman Puchkovskiy created ASF GitHub Bot updated a link from