Difference between revisions of "NutchWAX"

From COPTR
Jump to navigation Jump to search
(Import from spreadsheet via script.)
Line 1: Line 1:
{{Infobox_tool
+
{{Infobox tool
 
|purpose=NutchWAX is software for indexing ARC files (archived Web sites gathered using Heritrix) for full text search.
 
|purpose=NutchWAX is software for indexing ARC files (archived Web sites gathered using Heritrix) for full text search.
|image=
 
 
|homepage=http://archive-access.sourceforge.net/projects/nutchwax/
 
|homepage=http://archive-access.sourceforge.net/projects/nutchwax/
 
|license=GNU Lesser General Public License 2.1; Nutch itself is under Apache License 2.0.
 
|license=GNU Lesser General Public License 2.1; Nutch itself is under Apache License 2.0.
 
|platforms=Platform-independent Java, though only tested and primarily used on Linux machines.
 
|platforms=Platform-independent Java, though only tested and primarily used on Linux machines.
 +
|function=Web Crawl
 +
|content=Web
 +
}}
 +
{{Infobox tool details
 +
|ohloh_id=NutchWAX
 
}}
 
}}
 
<!-- Delete the Categories that do not apply -->
 
[[Category:Web Crawl]]
 
[[Category:Web]]
 
 
 
 
= Description =
 
= Description =
 
NutchWAX is software for indexing ARC files (archived Web sites gathered using Heritrix) for full text search. NutchWAX is based on the open-source Web-search software, Nutch. Developed by Internet Archive. Written in Java.
 
NutchWAX is software for indexing ARC files (archived Web sites gathered using Heritrix) for full text search. NutchWAX is based on the open-source Web-search software, Nutch. Developed by Internet Archive. Written in Java.
Line 19: Line 17:
  
 
= Development Activity =
 
= Development Activity =
 
{{Infobox_tool_details
 
|ohloh_id=NutchWAX
 
}}
 

Revision as of 16:23, 22 April 2021



NutchWAX is software for indexing ARC files (archived Web sites gathered using Heritrix) for full text search.
Homepage:http://archive-access.sourceforge.net/projects/nutchwax/
License:GNU Lesser General Public License 2.1; Nutch itself is under Apache License 2.0.
Platforms:Platform-independent Java, though only tested and primarily used on Linux machines.
Function:Web Crawl
Content type:Web


Error in widget Ohloh Project: unable to write file /var/www/html/extensions/Widgets/compiled_templates/wrt673fb0d8e0d414_02637023


Description

NutchWAX is software for indexing ARC files (archived Web sites gathered using Heritrix) for full text search. NutchWAX is based on the open-source Web-search software, Nutch. Developed by Internet Archive. Written in Java.

User Experiences

Development Activity