ParseKeggGenomeManager
extends ParseKeggAbstractManager
in package
Class ParseKeggGenomeManager A genome record names one sequenced organism KEGG holds pathways for, and ties it to the NCBI taxonomy. Not to be confused with ParseGenomeManager, which reads the sequencing statistics of the Legacy "DOGS" records.
Tags
Table of Contents
Properties
- $entry : string
- $names : array<string|int, mixed>
- $definition : string
- $lineage : array<string|int, mixed>
- $taxonomy : string
- NCBI taxonomy identifier, read out of the "TAX:" prefix the field writes it with.
Methods
- __construct() : mixed
- Constructor.
- getDefinition() : string
- getEntry() : string
- getEntryId() : string
- Extracts the identifier uniquely naming a KEGG record.
- getFormat() : string
- The name this format is known by in the collection records and in DatabaseParserFactory.
- getLineage() : array<string|int, mixed>
- getNames() : array<string|int, mixed>
- getTaxonomy() : string
- isEntryEnd() : bool
- Tells whether a line closes a KEGG record. KEGG closes on three slashes where most flat files use two.
- isEntryStart() : bool
- Tells whether a line opens a new KEGG record.
- parseDataFile() : Sequence
- Parses a KEGG genome data file and populates this manager's fields.
- joinLines() : string
- Joins the lines of a field into one string. A line opening with "$" continues the word the line above broke off, so it joins without a space.
- parseDbLinks() : array<string|int, mixed>
- Reads a DBLINKS field into pairs of database name and identifier, one per line.
- parsePathways() : array<string|int, mixed>
- Reads a PATHWAY field into pairs of map identifier and pathway name.
- readData() : string
- Reads the data a line carries, which starts at its thirteenth column.
- readEntryId() : string
- Reads the identifier out of an ENTRY line. Its last word names the kind of record rather than the record itself - "C00031 Compound", "EC 2.7.1.1 Enzyme" - so the identifier is what comes before it, which for an enzyme is the two words "EC" and its number.
- readFields() : array<string|int, mixed>
- Gathers a record into its fields : one entry per label, holding the data of its own line and of every continuation line below it.
- readLabel() : string
- Reads the label a line carries in its first twelve columns.
- splitTokens() : array<string|int, mixed>
- Splits the lines of a field into the whitespace-separated identifiers they list.
- readTaxonomy() : string
- Reads the TAXONOMY line, which carries the identifier behind a "TAX:" prefix.
Properties
$entry
protected
string
$entry
= ""
$names
protected
array<string|int, mixed>
$names
= []
$definition
private
string
$definition
= ""
$lineage
private
array<string|int, mixed>
$lineage
= []
$taxonomy
NCBI taxonomy identifier, read out of the "TAX:" prefix the field writes it with.
private
string
$taxonomy
= ""
Methods
__construct()
Constructor.
public
__construct() : mixed
getDefinition()
public
getDefinition() : string
Return values
stringgetEntry()
public
getEntry() : string
Return values
stringgetEntryId()
Extracts the identifier uniquely naming a KEGG record.
public
static getEntryId(array<string|int, mixed> $aFlines, string $sLine) : string
Parameters
- $aFlines : array<string|int, mixed>
-
The whole file, buffered
- $sLine : string
-
The line opening the entry
Return values
stringgetFormat()
The name this format is known by in the collection records and in DatabaseParserFactory.
public
static getFormat() : string
Return values
stringgetLineage()
public
getLineage() : array<string|int, mixed>
Return values
array<string|int, mixed>getNames()
public
getNames() : array<string|int, mixed>
Return values
array<string|int, mixed>getTaxonomy()
public
getTaxonomy() : string
Return values
stringisEntryEnd()
Tells whether a line closes a KEGG record. KEGG closes on three slashes where most flat files use two.
public
static isEntryEnd(string $sLine) : bool
Parameters
- $sLine : string
-
The line to analyze
Return values
boolisEntryStart()
Tells whether a line opens a new KEGG record.
public
static isEntryStart(string $sLine) : bool
Parameters
- $sLine : string
-
The line to analyze
Return values
boolparseDataFile()
Parses a KEGG genome data file and populates this manager's fields.
public
parseDataFile(array<string|int, mixed> $aFlines) : Sequence
Parameters
- $aFlines : array<string|int, mixed>
-
The lines the script has to parse
Tags
Return values
Sequence —$oSequence
joinLines()
Joins the lines of a field into one string. A line opening with "$" continues the word the line above broke off, so it joins without a space.
protected
joinLines(array<string|int, mixed> $aLines) : string
Parameters
- $aLines : array<string|int, mixed>
Return values
stringparseDbLinks()
Reads a DBLINKS field into pairs of database name and identifier, one per line.
protected
parseDbLinks(array<string|int, mixed> $aLines) : array<string|int, mixed>
Format : DBLINKS CAS: 50-99-7
Parameters
- $aLines : array<string|int, mixed>
Return values
array<string|int, mixed>parsePathways()
Reads a PATHWAY field into pairs of map identifier and pathway name.
protected
parsePathways(array<string|int, mixed> $aLines) : array<string|int, mixed>
Format : PATHWAY PATH: map00010 Glycolysis / Gluconeogenesis
Parameters
- $aLines : array<string|int, mixed>
Return values
array<string|int, mixed>readData()
Reads the data a line carries, which starts at its thirteenth column.
protected
static readData(string $sLine) : string
Parameters
- $sLine : string
-
The line to analyze
Return values
stringreadEntryId()
Reads the identifier out of an ENTRY line. Its last word names the kind of record rather than the record itself - "C00031 Compound", "EC 2.7.1.1 Enzyme" - so the identifier is what comes before it, which for an enzyme is the two words "EC" and its number.
protected
static readEntryId(string $sData) : string
Parameters
- $sData : string
Return values
stringreadFields()
Gathers a record into its fields : one entry per label, holding the data of its own line and of every continuation line below it.
protected
readFields(array<string|int, mixed> $aFlines) : array<string|int, mixed>
Parameters
- $aFlines : array<string|int, mixed>
-
The lines the script has to parse
Return values
array<string|int, mixed>readLabel()
Reads the label a line carries in its first twelve columns.
protected
static readLabel(string $sLine) : string
Parameters
- $sLine : string
-
The line to analyze
Return values
stringsplitTokens()
Splits the lines of a field into the whitespace-separated identifiers they list.
protected
splitTokens(array<string|int, mixed> $aLines) : array<string|int, mixed>
Parameters
- $aLines : array<string|int, mixed>
Return values
array<string|int, mixed>readTaxonomy()
Reads the TAXONOMY line, which carries the identifier behind a "TAX:" prefix.
private
readTaxonomy(string $sData) : string
Format : TAXONOMY TAX:9606
Parameters
- $sData : string