Greenstone Wiki

ReadTextFile plugin

Base plugin for files that are plain text.

The following table lists all of the configuration options available for ReadTextFile.

Option	Description	Value
ReadTextFile Options
input_encoding	The encoding of the source documents. Documents will be converted from these encodings and stored internally as utf8.	Default: auto List
default_encoding	Use this encoding if -input_encoding is set to 'auto' and the text categorization algorithm fails to extract the encoding or extracts an encoding unsupported by Greenstone. This option can take the same values as -input_encoding.	Default: utf8 List
extract_language	Identify the language of each document and set 'Language' metadata. Note that this will be done automatically if -input_encoding is 'auto'.
default_language	If Greenstone fails to work out what language a document is the 'Language' metadata element will be set to this value. The default is 'en' (ISO 639 language symbols are used: en = English). Note that if -input_encoding is not set to 'auto' and -extract_language is not set, all documents will have their 'Language' metadata set to this value.	Default: en
Options Inherited from AutoExtractMetadata
first	Comma separated list of numbers of characters to extract from the start of the text into a set of metadata fields called 'FirstN', where N is the size. For example, the values "3,5,7" will extract the first 3, 5 and 7 characters into metadata fields called "First3", "First5" and "First7".
Options Inherited from AcronymExtractor
extract_acronyms	Extract acronyms from within text and set as metadata.
markup_acronyms	Add acronym metadata into document text.
Options Inherited from KeyphraseExtractor
extract_keyphrases	Extract keyphrases automatically with Kea (default settings).
extract_keyphrases_kea4	Extract keyphrases automatically with Kea 4.0 (default settings). Kea 4.0 is a new version of Kea that has been developed for controlled indexing of documents in the domain of agriculture.
extract_keyphrase_options	Options for keyphrase extraction with Kea. For example: mALIWEB - use ALIWEB extraction model; n5 - extract 5 keyphrase;, eGBK - use GBK encoding.
Options Inherited from EmailAddressExtractor
extract_email	Extract email addresses as metadata.
Options Inherited from DateExtractor
extract_historical_years	Extract time-period information from historical documents. This is stored as metadata with the document. There is a search interface for this metadata, which you can include in your collection by adding the statement, "format QueryInterface DateSearch" to your collection configuration file.
maximum_year	The maximum historical date to be used as metadata (in a Common Era date, such as 1950).	Default: 2013
maximum_century	The maximum named century to be extracted as historical metadata (e.g. 14 will extract all references up to the 14th century).	Default: -1
no_bibliography	Do not try to block bibliographic dates when extracting historical dates.
Options Inherited from GISExtractor
extract_placenames	Extract placenames from within text and set as metadata. Requires GIS extension to Greenstone.
gazetteer	Gazetteer to use to extract placenames from within text and set as metadata. Requires GIS extension to Greenstone.
place_list	When extracting placements, include list of placenames at start of the document. Requires GIS extension to Greenstone.
Options Inherited from BasePlugin
process_exp	A perl regular expression to match against filenames. Matching filenames will be processed by this plugin. For example, using '(?i).html?\$' matches all documents ending in .htm or .html (case-insensitive).
no_blocking	Don't do any file blocking. Any associated files (e.g. images in a web page) will be added to the collection as documents in their own right.
block_exp	Files matching this regular expression will be blocked from being passed to any later plugins in the list.
store_original_file	Save the original source document as an associated file. Note this is already done for files like PDF, Word etc. This option is only useful for plugins that don't already store a copy of the original file.
associate_ext	Causes files with the same root filename as the document being processed by the plugin AND a filename extension from the comma separated list provided by this argument to be associated with the document being processed rather than handled as a separate list.
associate_tail_re	A regular expression to match filenames against to find associated files. Used as a more powerful alternative to associate_ext.
OIDtype	The method to use when generating unique identifiers for each document.	Default: auto List
OIDmetadata	Specifies the metadata element that hold's the document's unique identifier, for use with -OIDtype=assigned.	Default: dc.Identifier
no_cover_image	Do not look for a prefix.jpg file (where prefix is the same prefix as the file being processed) to associate as a cover image.
filename_encoding	The encoding of the source file filenames.	Default: auto List
file_rename_method	The method to be used in renaming the copy of the imported file and associated files.	Default: url List

input_encoding option values

Value	Description
auto	Use text categorization algorithm to automatically identify the encoding of each source document. This will be slower than explicitly setting the encoding but will work where more than one encoding is used within the same collection.
ascii	Plain 7 bit ASCII. This may be a bit faster than using iso_8859_1. Beware of using this when the text may contain characters outside the plain 7 bit ASCII set though (e.g. German or French text containing accents), use iso_8859_1 instead.
utf8	Either utf8 or unicode – automatically detected.
unicode	Just unicode.
iso_8859_6	Arabic
gb	Chinese Simplified (GB)
big5	Chinese Traditional (Big5)
koi8_r	Cyrillic
iso_8859_5	Cyrillic
koi8_u	Cyrillic (Ukrainian)
dos_437	DOS codepage 437 (US English)
dos_850	DOS codepage 850 (Latin 1)
dos_852	DOS codepage 852 (Central European)
dos_866	DOS codepage 866 (Cyrillic)
iso_8859_7	Greek
iso_8859_8	Hebrew
iscii_de	ISCII Devanagari
euc_jp	Japanese (EUC)
shift_jis	Japanese (Shift-JIS)
korean	Korean (Unified Hangul Code - i.e. a superset of EUC-KR)
iso_8859_1	Latin1 (western languages)
iso_8859_15	Latin15 (revised western)
iso_8859_2	Latin2 (central and eastern european languages)
iso_8859_3	Latin3
iso_8859_4	Latin4
iso_8859_9	Turkish
windows_1250	Windows codepage 1250 (WinLatin2)
windows_1251	Windows codepage 1251 (WinCyrillic)
windows_1252	Windows codepage 1252 (WinLatin1)
windows_1253	Windows codepage 1253 (WinGreek)
windows_1254	Windows codepage 1254 (WinTurkish)
windows_1255	Windows codepage 1255 (WinHebrew)
windows_1256	Windows codepage 1256 (WinArabic)
windows_1257	Windows codepage 1257 (WinBaltic)
windows_1258	Windows codepage 1258 (Vietnamese)
windows_874	Windows codepage 874 (Thai)

default_encoding option values

Value	Description
ascii	Plain 7 bit ASCII. This may be a bit faster than using iso_8859_1. Beware of using this when the text may contain characters outside the plain 7 bit ASCII set though (e.g. German or French text containing accents), use iso_8859_1 instead.
utf8	Either utf8 or unicode – automatically detected.
unicode	Just unicode.
iso_8859_6	Arabic
gb	Chinese Simplified (GB)
big5	Chinese Traditional (Big5)
koi8_r	Cyrillic
iso_8859_5	Cyrillic
koi8_u	Cyrillic (Ukrainian)
dos_437	DOS codepage 437 (US English)
dos_850	DOS codepage 850 (Latin 1)
dos_852	DOS codepage 852 (Central European)
dos_866	DOS codepage 866 (Cyrillic)
iso_8859_7	Greek
iso_8859_8	Hebrew
iscii_de	ISCII Devanagari
euc_jp	Japanese (EUC)
shift_jis	Japanese (Shift-JIS)
korean	Korean (Unified Hangul Code - i.e. a superset of EUC-KR)
iso_8859_1	Latin1 (western languages)
iso_8859_15	Latin15 (revised western)
iso_8859_2	Latin2 (central and eastern european languages)
iso_8859_3	Latin3
iso_8859_4	Latin4
iso_8859_9	Turkish
windows_1250	Windows codepage 1250 (WinLatin2)
windows_1251	Windows codepage 1251 (WinCyrillic)
windows_1252	Windows codepage 1252 (WinLatin1)
windows_1253	Windows codepage 1253 (WinGreek)
windows_1254	Windows codepage 1254 (WinTurkish)
windows_1255	Windows codepage 1255 (WinHebrew)
windows_1256	Windows codepage 1256 (WinArabic)
windows_1257	Windows codepage 1257 (WinBaltic)
windows_1258	Windows codepage 1258 (Vietnamese)
windows_874	Windows codepage 874 (Thai)

Table of Contents

ReadTextFile plugin

input_encoding option values

default_encoding option values