Ling Links

health: probiotics and mood

family tree software

Grep Regex



grep  "twitter \|wikipedia " filename                        # grep for A or B
grep -P  "twitter |wikipedia " filename                     # grep for A or B
grep "\\\\" filename                                               # grep for a single \



grep -P "\t" filename                                             # grep for tab

grep -h          - suppress filename content is found in
grep -A 2       - returns matching line and to lines After
grep -B 2       - returns matching line and to lines Before


Grep Regular Expressions

grep regex
http://www.bo.infn.it/alice/alice-doc/mll-doc/usrgde/node28.html

more about unicode hex etc

common python snippets



Python Modules  http://wiki.python.org/moin/UsefulModules


regex subsitution in myString
myString= re.sub("\([0-9][0-9]*\)", "", myString)
myString= re.sub("\\\\", "", myString)      # remove a single backslash
myString= re.sub(r"\\", "", myString)      # remove a single backslash using a "raw string"


string substitution (*not* for regex)
    c = c.replace(u'\xae', '') # ®
    c = c.replace(u'\xbb', '') # »
    c = c.replace(u'\x99', '') # char for TM
    c = c.replace(u'\xa9', '') # ©
    c = c.replace(u'—', '-')  # replace mdash u2014 with regular dash
    c = c.replace('?', '') # ?
    c = c.replace(':', '') # :
    c = c.replace(';', '') # ;




save regex-matching string to a new variable
        pageObj = re.search("od [0-9][0-9]*", pagetext)
        if pageObj:
            pageObj2 = re.search("[0-9][0-9]*", pageObj.group())
            numb = pageObj2.group()
            print "numb is ",numb
            pages = range(2,int(numb)+1)

            for p in pages:
                newUrl = locUrl+'afpg/'+str(p)+'/Default.aspx'
                print "adding ", newUrl

stuff to write about



BIO notation

Inside-Outside Algorithm

HPSG

HPSG (Head-Driven Phrase Structure Grammar) 

Stanford and Berkeley espouse this syntactic theory.
It is a "highly lexicalized, non-derivational generative grammar theory."
It is a type of phrase structure grammar, as oppsed to a dependency grammar (The syntactic trees I learned are dependency grammars).


sed one-liners


sed 's/ Телефон: /\t/' filename             # basic search and replace
sed 's/ [-] .*//' filename                       # put tricky chars in []
sed 's/\&amp/\&/' filename                  # example of escaped char
sed 's/.\{1\}/& /g' filename                   # insert 1 space between each char
sed 's/.\{2\}/& /g' filename                   # insert 2 spaces between each char
sed 's/\(.*\)\/word$/word\1/' filename     # search and replace with a captured pattern () on LHS and \1 on RHS
sed 's/word/&\n/' filename                   # search and replace insert a newline
sed '/^\s*$/d' filename                         # delete blank line




Good resources for one-liners with examples 
http://sed.sourceforge.net/sed1line.txt

Starting Off in Python


We’ll start off by looking at material from these sites.
http://www.astro.ufl.edu/~warner/prog/python.html
http://www.sthurlow.com/python/

http://learnpythonthehardway.org/book/ <- recommended for new programmers

using "screen" command to handle multiple sessions

from http://www.cyberciti.biz/tips/linux-screen-command-howto.html


$ screen -S 1
CTRL+a, c            -- create another screen window
CTRL+a, n            -- switch to next screen window I've got open




  • To list all windows use the command CTRL+a followed by " key (first hit CTRL+a, releases both keys and press " ).
  • To switch to window by number use the command CTRL+a followed by ' (first hit CTRL+a, releases both keys and press ' it will prompt for window number).

Common screen commands

screen commandTask
Ctrl+a cCreate new window
Ctrl+a kKill the current window / session
Ctrl+a wList all windows
Ctrl+a 0-9Go to a window numbered 0 9, use Ctrl+a w to see number
Ctrl+a Ctrl+aToggle / switch between the current and previous window
Ctrl+a SSplit terminal horizontally into regions and press Ctrl+a c to create new window there
Ctrl+a :resizeResize region
Ctrl+a :fitFit screen size to new terminal size. You can also hit Ctrl+a F for the the same task
Ctrl+a :removeRemove / delete region. You can also hit Ctrl+a X for the same taks
Ctrl+a tabMove to next region
Ctrl+a D (Shift-d)Power detach and logout
Ctrl+a dDetach but keep shell window open
Ctrl-a Ctrl-\Quit screen
Ctrl-a ?Display help screen i.e. display a list of commands

Suggested readings:

See screen command man page for further details:
man screen


awk

specify tab-delimited columns
awk -F'\t' '{ print $1}' file

print column $13 and column $10 where $13 matches "MY"
awk -F"\t" '$13 ~ /MY/ {print $13"\t"$10}' | less

awk -F"\t" '$7 == "restaurant" ' japan.osm.poi  > restaurants


awk '{ print \$1} ' file,   print 1st column, escape $ when in perl system command



not really sure what this does.....

awk -F "\"*,\"*" '{print $3"\t"$5}' jigyosyo.csv

iconv and recode to change file encoding


List all the encoding codes:
iconv --list

iconv --from-code LATIN1 --to-code UTF-8 --output adoos.com.my.categories.zlm-MYS.UTF.txt adoos.com.my.categories.zlm-MYS.LATIN1.txt

recode ....

windows 1252 encoding

CP1252 is windows encoding

em dash is a measurement of font size that is often double encoded?
to find it do:
zcat file | grep -P '\xc2\x96'

91,92,93,94 are also other troublesome windows chars




my .bashrc file


So you don't have to do source ~/.bashrc every time you open a terminal, put it in ~/.bash_profile.
~> cat .bash_profile
source ~/.bashrc


Things I like to have in my .bashrc file are:
.... in progress

file transfer

If you're transferring between Windows and linux, use winSCP.

If you're transferring linux to linux use scp like this:

>> scp ......

EM (Expectation maximization) algorithm

- an iterative method for find maximum likelihood estimates of parameters in a statistical model
   - E compute the expectation of the log-likelihood using the current estimates for the parameters
  - M compute parameters maximizing the expected log-likelihood found during the E step




I was trying to find weights for rules of my PCFG.
I used gigaword news data.

Comp Ling Glossary


  • context free grammar (CFG) - in formal language theory a formal grammar where every production rule has the form A-> b where A is non-terminal and b is 0+ terminal and/or nonterminals.
    • linguistics sometimes calls them phrase structure grammars
    • comp sci often uses Backus-Naur Form (BNF) for CFGs.
      <symbol> ::= __expression__
      <US-address> ::= <name> <street> <zip> "USA"
  • perplexity - a measure of how well test data is predicted by a model.  
    • I read that the lowest perplexity found using a 3gram on the Brown Corpus is 247.
    • The perfect/true model for the data would have a perplexity of 0.
  • precision & recall - In IR precision is the fraction of retrieved instances that are relevant and recall is the fraction of relevant instances that are retrieved.
    • Max precision is no false positives (being conservative about finding a match).
    • Max recall is no false negatives (being exhaustive and thorough).

Morphology review, glossary


(less complex morphology                              more complex morphology)
(low morpheme-to-word ratio                           high morpheme-to-word ratio)
analytic / isolating > agglutinative / fusional > polysynthetic
  • lexeme – a word concept.  A lexeme can represent multiple surface forms of a word that comprise the paradigm of the lexeme (e.g. run, runs, ran and running are the same lexeme, RUN).  As parts of a single concept, the forms of a lexeme will have the same syntactic category (runner is a different lexeme).
  • lemma – a particular form of a lexeme conventionally chosen to represent the canonical form.  In a dictionary the lemma is the headword.
  • paradigm - complete set of words associate with a given lexeme.
  • inflectional rules – relate a lexeme to its forms.  Generally the syntactic category remains the same (e.g. eat and eaten, boy and boys)
  • derivational rules – relate a lexeme to a new lexeme.  Generally, but not always, the syntactic category changes (e.g. slow and slowly).  Necessarily the meaning of the base changes (e.g. write and rewrite, circle and encircle).  Derivational affixes are bound morphemes.
  • zero derivation or conversion – changing from one lexeme to another without any surface change (e.g. telephone and to telephone).

  • endocentric compound – The compound has a head which represents the basic meaning (and same part-of-speech) of the whole (e.g. house and doghouse)
  • exocentric compound – meaning of the compound is not transparent from the constituents (e.g. white-collar and must-have).
  • copulative – …
  • appositional – …

Review of common grammatical cases

 

“Among modern languages, cases still feature prominently in most of the Balto-Slavic languages, with most having six to eight cases, as well as German and Modern Greek, which have four.  In German, cases are mostly market on articles and adjectives, and less so on nouns.” (wikipedia)

  1. nominative – subject of finite verb
  2. accusative – direct object of verb
  3. dative – indirect object of a verb
  4. ablative – indicates movement from something or causality
  5. genitive – possessive
  6. vocative – addressee
  7. locative – a location
  8. instrumental – object use in performing an action