regex-base-0.72.0.2: doc/lazy.html
<!DOCTYPE HTML PUBLIC "-//W3C//DTD HTML 4.0 Transitional//EN">
<HTML>
<head>
<title>Text.Regex.Lazy</title>
</head>
<body>
<h1><tt>Text.Regex.Lazy</tt></h1>
<h2>Version 0.70 (2006-08-10)</h2>
<h3>By Chris Kuklewicz (TextRegexLazy (at) personal (dot) mightyreason (dot) com)</h3>
Changes from 0.66 to 0.70
<ul>
<li> regex-tre added for libtre backend (Text.Regex.TRE), see http://laurikari.net/tre/
<li> regex-devel added for tests and benchmarks
<li> Text.Regex.*.Wrap APIs improved: the exported wrap* functions
never call fail or error under normal circumstances, and use Either
types to report errors. Allocation failures are reported with fail.
<li> Text.Regex.*.(ByteString|String) all should export
compile/execute/regexec functions which report errors using Either.
</ul>
Changes from 0.55 to 0.66
<ul>
<li> I broke this into many packages, regex-base for the interface and regex-pcre, regex-posix, regex-parsec, regex-dfa for the four backends and regex-compat to replace Text.Regex(.New)
<li> The top level Makefile now can drive setup and installation of all the packages at once.
</ul>
Changes from 0.44 to 0.55
<ul>
<li> <b>JRegex has been assimilated: PCRE and PosixRE are here</b>.
The JRegex-style API rocks, see below and Context.hs and Example.hs
<li> Haddock seems to run via ./setup haddock, but the documentation is very thin
<li> ./setup test runs TestTestRegexLazy binary if uncommented in cabal file
<li> default is now to compile with -Wall -Werror -O2
<li> You may need to point the cabal file's "Extra-Lib-Dirs" to point to pcre.
<li> You may or may not need a "-lpcre" option to ghc when building
projects that depend on Text.Regex.Lazy now.
</ul>
Changes from 0.33 to 0.44
<ul>
<li> Cabal
<li> Compile with -Wall -Werror
<li> Change DFAEngineFPS from Data.FastPackedString to Data.ByteString
</ul>
See the LICENSE file for details on copyright. See README for building instructions.
<br/>
The new API is very close to JRegex and supports 4 backends:
<ul>
<li> Posix, the standard c regex library
<li> PCRE, the <a href="http://www.pcre.org/">Perl Compatible Regular Expressions</a> c library
<li> Full, the lazy Parsec based library (see old api below)
<li> DFA, the fast lazy matching library (see old api below)
</ul>
And for all backends, there are two types that can be used as a source
of regular expressions or to match a regular expression against:
String, and ByteString. The ByteString library will be in the next
GHC and can be gotten
from <a
href="http://www.cse.unsw.edu.au/~dons/fps.html">http://www.cse.unsw.edu.au/~dons/fps.html</a>.
<p>
For simplest use of the new API: import Text.Regex.Lazy and one of
<pre>
import Text.Regex.PCRE((=~),(=~~))
import Text.Regex.Parsec((=~),(=~~))
import Text.Regex.DFA((=~),(=~~))
import Text.Regex.PosixRE((=~),(=~~))
import Text.Regex.TRE((=~),(=~~))
</pre>
The things you can demand of (=~) and (=~~) are all
instance defined in Text.Regex.Impl.Context and they are used
in <tt>Example.hs</tt> as well.
<p>
<p>
You can redefine (=~) and (=~~) to use different options by using makeRegexOpts:
<pre>
(=~) :: (RegexMaker Regex CompOption ExecOption source,RegexContext Regex source1 target) => source1 -> source -> target
(=~) x r = let q :: Regex
q = makeRegexOpts (some compoption) (some execoption) r
in match q x
(=~~) ::(RegexMaker Regex CompOption ExecOption source,RegexContext Regex source1 target,Monad m) => source1 -> source -> m target
(=~~) x r = let q :: Regex
q = makeRegexOpts (some compoption) (some execoption) r
in matchM q x
</pre>
There is a medium level API with functions compile/execute/regexec in
all the Text.Regex.*.(String|ByteString) modules. These allow for
errors to be reported as Either types when compiling or running.
<p>
The low level APIs are in the Text.Regex.*.Wrap modules. For the
c-library backends these expose most of the c-api in wrap* functions
that make the type more Haskell-like: CString and CStingLen and
newtypes to specify compile and execute options. The actual foreign
calls are not exported; it does not export the raw c api.
<p>
Also, Text.Regex.PCRE.Wrap will let you query if it was compiled with
UTF8 suppor: <tt>configUTF8 :: Bool</tt>. But I do not provide a way
to marshall to or from UTF8. (If you have a UTF8 ByteString then you
would probably be able to make it work, assuming the indices PCRE uses
are in bytes, otherwise look at the wrap* functions which are a thin
layer over the pcreapi).
<p>
<p>
The old Text.Regex API is can be replaced. If you need to be drop in
compatible with <tt>Text.Regex</tt> then you can
import <tt>Text.Regex.New</tt> and report any infidelities as bugs.
Some advantages of <tt>Text.Regex.Parsec</tt> over <tt>Text.Regex</tt>:
<ul>
<li> It does not marshal to and from c-code arrays, so it is much
faster on large input strings.
<li> It consumes the input <tt>String</tt> in a mostly lazy manner.
This makes streaming from input to output possible.
<li> It performs sanity checks so that <tt>subRegex</tt>
and <tt>splitRegex</tt> don't loop or go crazy if the pattern
matches an empty string -- it will just return the input.
<li> If the <tt>String</tt> regex does not parse then you get a nicer error
message.
</ul>
<p>
Internally it uses <tt>Parsec</tt> to turn the string regex into
a <tt>Pattern</tt> data type, simplify the <tt>Pattern</tt>, then
transform the <tt>Pattern</tt> into a <tt>Parsec</tt> parser that
accepts matching strings and stores the sub-strings of parenthesized
groups.
<p>
All of this was motivated by the inability to use <tt>Text.Regex</tt>
to complete
the <a
href="http://shootout.alioth.debian.org/gp4/benchmark.php?test=regexdna&lang=all">regex-dna
benchmark</a> on <a href="http://shootout.alioth.debian.org/">The
Computer Language Shootout</a>. The current entry there, by Don
Stewart and Alson Kemp and Chris Kuklewicz, does not use this Parsec
solution, but rather a custom DFA lexer from the CTK library.
</body>
</HTML>