replace-attoparsec 1.2.0.0 → 1.2.1.0
raw patch · 6 files changed
+91/−60 lines, 6 filesPVP ok
version bump matches the API change (PVP)
API changes (from Hackage documentation)
Files
- README.md +24/−19
- replace-attoparsec.cabal +1/−1
- src/Replace/Attoparsec/ByteString.hs +30/−19
- src/Replace/Attoparsec/Text.hs +31/−18
- tests/TestByteString.hs +3/−1
- tests/TestText.hs +2/−2
README.md view
@@ -63,6 +63,7 @@ the regular expression says it has found a numeric string but the string-to-number conversion function fails. A typed parser will perform both the pattern match and the conversion, so it will never be in that situation.+ [Parse, don't validate.](https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-validate/) * Regular expressions are only able to pattern-match [regular](https://en.wikipedia.org/wiki/Chomsky_hierarchy#The_hierarchy)@@ -227,7 +228,7 @@ #!/usr/bin/env stack {- stack script- --resolver nightly-2019-09-13+ --resolver lts-15 --package attoparsec --package text --package text-show@@ -322,27 +323,28 @@ | Program | dense | sparse | | :--- | ---: | ---: |-| Python `re.sub`¹ | 89.23ms | 23.98ms |-| Perl `s///ge`² | 180.65ms | 5.60ms |-| [`Replace.Megaparsec.streamEdit`][m] `String` | 454.95ms | 375.04ms |+| [Python 3.7.4 `re.sub`][sub] *repl* function | 89.23ms | 23.98ms |+| [Perl 5 `s///ge`][s] | 180.65ms | 5.02ms |+| [`Replace.Megaparsec.streamEdit`][m] `String` | 441.94ms | 375.04ms | | [`Replace.Megaparsec.streamEdit`][m] `ByteString` | 529.99ms | 73.76ms | | [`Replace.Megaparsec.streamEdit`][m] `Text` | 547.47ms | 139.21ms | | [`Replace.Attoparsec.ByteString.streamEdit`][ab] | 394.12ms | 41.13ms | | [`Replace.Attoparsec.Text.streamEdit`][at] | 515.26ms | 46.10ms | | [`Text.Regex.Applicative.replace`][ra] `String` | 1083.98ms | 646.40ms |-| [`Text.Regex.PCRE.Heavy.gsub`][ph] `Text` | ⊥³ | 14.76ms |--¹ Python 3.7.4--² This is perl 5, version 28, subversion 2 (v5.28.2) built for x86_64-linux-thread-multi--³ Does not finish.+| [`Text.Regex.PCRE.Heavy.gsub`][ph] `Text` | > 10min | 14.29ms |+| [`Control.Lens.Regex.ByteString.match`][lb] | > 10min | 4.27ms |+| [`Control.Lens.Regex.Text.match`][lt] | > 10min | 14.74ms | +[sub]: https://docs.python.org/3/library/re.html#re.sub+[s]: https://perldoc.perl.org/functions/s.html [m]: https://hackage.haskell.org/package/replace-megaparsec/docs/Replace-Megaparsec.html#v:streamEdit [ab]: https://hackage.haskell.org/package/replace-attoparsec/docs/Replace-Attoparsec-ByteString.html#v:streamEdit [at]: https://hackage.haskell.org/package/replace-attoparsec/docs/Replace-Attoparsec-Text.html#v:streamEdit [ra]: http://hackage.haskell.org/package/regex-applicative/docs/Text-Regex-Applicative.html#v:replace+[ss]: http://hackage.haskell.org/package/stringsearch/docs/Data-ByteString-Search.html#v:replace [ph]: http://hackage.haskell.org/package/pcre-heavy/docs/Text-Regex-PCRE-Heavy.html#v:gsub+[lb]: https://hackage.haskell.org/package/lens-regex-pcre/docs/Control-Lens-Regex-ByteString.html#v:match+[lt]: https://hackage.haskell.org/package/lens-regex-pcre/docs/Control-Lens-Regex-Text.html#v:match # Hypothetically Asked Questions@@ -356,19 +358,22 @@ 2. *Is this a good idea?* - You may have heard it suggested that monadic parsers are better when+ You may have+ [heard it suggested](https://stackoverflow.com/questions/57667534/how-can-i-use-a-parser-in-haskell-to-find-the-locations-of-some-substrings-in-a/57712672#comment101804063_57667534)+ that monadic parsers are better for pattern-matching when the input stream is mostly signal, and regular expressions are better when the input stream is mostly noise. - The premise of this library is:- that sentiment is outdated; monadic parsers are great for finding- small patterns in a stream of otherwise uninteresting text; and the- reluctance to forego the speedup opportunities afforded by restricting+ The premise of this library is that monadic parsers are great for finding+ small signal patterns in a stream of otherwise noisy text.++ Our reluctance to forego the speedup opportunities afforded by restricting ourselves to regular grammars is an old superstition about opportunities which [remain mostly unexploited anyway](https://swtch.com/~rsc/regexp/regexp1.html).- The performance compromise of allowing stack memory allocation (a.k.a pushdown- automata, a.k.a context-free grammar) was once considered- [controversial for *general-purpose* programming languages](https://vanemden.wordpress.com/2014/06/18/how-recursion-got-into-programming-a-comedy-of-errors-3/). I think we+ The performance compromise of allowing stack memory allocation (a.k.a. pushdown+ automata, a.k.a. context-free grammar) was once considered+ [controversial for *general-purpose* programming languages](https://vanemden.wordpress.com/2014/06/18/how-recursion-got-into-programming-a-comedy-of-errors-3/).+ I think we can now resolve that controversy the same way for pattern matching languages.
replace-attoparsec.cabal view
@@ -1,5 +1,5 @@ name: replace-attoparsec-version: 1.2.0.0+version: 1.2.1.0 cabal-version: 1.18 synopsis: Find, replace, and edit text patterns with Attoparsec parsers homepage: https://github.com/jamesdbrock/replace-attoparsec
src/Replace/Attoparsec/ByteString.hs view
@@ -58,34 +58,45 @@ -- == Separate and capture -- -- Parser combinator to find all of the non-overlapping ocurrences--- of the pattern @sep@ in a text stream. Separate the stream into sections:+-- of the pattern @sep@ in a text stream.+-- The 'sepCap' parser will always consume its entire input and can never fail. ----- * sections which can parsed by the pattern @sep@ will be captured as--- matching sections in 'Right'--- * non-matching sections of the stream will be captured in 'Left'.+-- === Output ----- This parser will always consume its entire input and can never fail.--- If there are no pattern matches, then the entire input stream will be--- returned as a non-matching 'Left' section.+-- The input stream is separated and output int a list of sections: ----- The pattern matching parser @sep@ will not be allowed to succeed without--- consuming any input. If we allow the parser to match a zero-width pattern,+-- * Sections which can parsed by the pattern @sep@ will be parsed and captured+-- as 'Right'+-- * Non-matching sections of the stream will be captured in 'Left'.+--+-- The output list also has these properties:+--+-- * If the input is @""@ then the output list will be @[]@.+-- * If there are no pattern matches, then+-- the entire input stream will be returned as one non-matching 'Left' section.+-- * The output list will not contain two consecutive 'Left' sections.+--+-- === Zero-width matches forbidden+--+-- If the pattern matching parser @sep@ would succeed without consuming any+-- input then 'sepCap' will force it to fail.+-- If we allow @sep@ to match a zero-width pattern, -- then it can match the same zero-width pattern again at the same position -- on the next iteration, which would result in an infinite number of--- overlapping pattern matches. So, for example, the--- pattern @many digit@, which can match zero occurences of a digit,--- will be treated by @sepCap@ as @many1 digit@, and required to match--- at least one digit.+-- overlapping pattern matches. --+-- === Notes+-- -- This @sepCap@ parser combinator is the basis for all of the other--- features of this module. It is similar to the @sep*@ family of functions--- found in+-- features of this module.+--+-- It is similar to the @sep*@ family of functions found in -- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html parser-combinators> -- and -- <http://hackage.haskell.org/package/parsers/docs/Text-Parser-Combinators.html parsers> -- but, importantly, it returns the parsed result of the @sep@ parser instead--- of throwing it away.---+-- of throwing it away, like+-- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html#v:manyTill_ manyTill_>. sepCap :: Parser a -- ^ The pattern matching parser @sep@ -> Parser [Either B.ByteString a]@@ -228,8 +239,8 @@ -- -- This allows us to write an @editor@ function which can choose to not -- edit the match and just leave it as it is. If the @editor@ function--- always returns the first item in the tuple, then @streamEdit@ changes--- nothing.+-- returns the first item in the tuple, then @streamEdit@ will not change+-- the matched string. -- -- So, for all @sep@: --
src/Replace/Attoparsec/Text.hs view
@@ -59,33 +59,46 @@ -- == Separate and capture -- -- Parser combinator to find all of the non-overlapping ocurrences--- of the pattern @sep@ in a text stream. Separate the stream into sections:+-- of the pattern @sep@ in a text stream.+-- The 'sepCap' parser will always consume its entire input and can never fail. ----- * sections which can parsed by the pattern @sep@ will be captured as--- matching sections in 'Right'--- * non-matching sections of the stream will be captured in 'Left'.+-- === Output ----- This parser will always consume its entire input and can never fail.--- If there are no pattern matches, then the entire input stream will be--- returned as a non-matching 'Left' section.+-- The input stream is separated and output int a list of sections: ----- The pattern matching parser @sep@ will not be allowed to succeed without--- consuming any input. If we allow the parser to match a zero-width pattern,+-- * Sections which can parsed by the pattern @sep@ will be parsed and captured+-- as 'Right'+-- * Non-matching sections of the stream will be captured in 'Left'.+--+-- The output list also has these properties:+--+-- * If the input is @""@ then the output list will be @[]@.+-- * If there are no pattern matches, then+-- the entire input stream will be returned as one non-matching 'Left' section.+-- * The output list will not contain two consecutive 'Left' sections.+--+-- === Zero-width matches forbidden+--+-- If the pattern matching parser @sep@ would succeed without consuming any+-- input then 'sepCap' will force it to fail.+-- If we allow @sep@ to match a zero-width pattern, -- then it can match the same zero-width pattern again at the same position -- on the next iteration, which would result in an infinite number of--- overlapping pattern matches. So, for example, the--- pattern @many digit@, which can match zero occurences of a digit,--- will be treated by @sepCap@ as @many1 digit@, and required to match--- at least one digit.+-- overlapping pattern matches. --+-- === Notes+-- -- This @sepCap@ parser combinator is the basis for all of the other--- features of this module. It is similar to the @sep*@ family of functions--- found in+-- features of this module.+--+-- It is similar to the @sep*@ family of functions found in -- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html parser-combinators> -- and -- <http://hackage.haskell.org/package/parsers/docs/Text-Parser-Combinators.html parsers> -- but, importantly, it returns the parsed result of the @sep@ parser instead--- of throwing it away.+-- of throwing it away, like+-- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html#v:manyTill_ manyTill_>.+ -- sepCap :: Parser a -- ^ The pattern matching parser @sep@@@ -247,8 +260,8 @@ -- -- This allows us to write an @editor@ function which can choose to not -- edit the match and just leave it as it is. If the @editor@ function--- always returns the first item in the tuple, then @streamEdit@ changes--- nothing.+-- returns the first item in the tuple, then @streamEdit@ will not change+-- the matched string. -- -- So, for all @sep@: --
tests/TestByteString.hs view
@@ -12,7 +12,6 @@ import "parsers" Text.Parser.Token import Replace.Attoparsec.ByteString import Control.Applicative-import Data.Monoid tests :: IO [Test] tests = return@@ -44,10 +43,13 @@ (sepCap (string "aa")) (" a") ("a ") ([Left " ",Right"aa",Left" "])+ , Test $ runParserTest "empty input" (sepCap (fail "" :: Parser ())) "" [] , Test $ streamEditTest "x to o" (string "x") (const "o") "x x x" "o o o" , Test $ streamEditTest "x to o inner" (string "x") (const "o") " x x x " " o o o " , Test $ streamEditTest "ordering" (string "456") (const "ABC") "123456789" "123ABC789"+ , Test $ streamEditTest "empty input" (match (fail "")) (fst) "" "" ]+ where runParserTest nam p input expected = TestInstance { run = do
tests/TestText.hs view
@@ -7,11 +7,9 @@ import Distribution.TestSuite as TestSuite import Replace.Attoparsec.Text import Data.Attoparsec.Text as A--- import Data.Attoparsec.Text.Char import qualified Data.Text as T import Text.Parser.Char (upper) import Control.Applicative-import Data.Monoid tests :: IO [Test] tests = return@@ -51,9 +49,11 @@ (findAll ((A.takeWhile (=='𝅘𝅥𝅯') :: Parser T.Text))) ("𝄞𝅘𝅥𝅘𝅥𝅘𝅥𝅘𝅥𝅘𝅥𝅯𝅘𝅥𝅯") ("𝅘𝅥𝅯𝅘𝅥𝅯𝅘𝅥𝅘𝅥𝅘𝅥𝅘𝅥" :: T.Text) [Left "𝄞𝅘𝅥𝅘𝅥𝅘𝅥𝅘𝅥", Right "𝅘𝅥𝅯𝅘𝅥𝅯𝅘𝅥𝅯𝅘𝅥𝅯", Left "𝅘𝅥𝅘𝅥𝅘𝅥𝅘𝅥"]+ , Test $ runParserTest "empty input" (sepCap (fail "" :: Parser ())) "" [] , Test $ streamEditTest "x to o" (string "x") (const "o") "x x x" "o o o" , Test $ streamEditTest "x to o inner" (string "x") (const "o") " x x x " " o o o " , Test $ streamEditTest "ordering" (string "456") (const "ABC") "123456789" "123ABC789"+ , Test $ streamEditTest "empty input" (match (fail "")) (fst) "" "" ] where runParserTest nam p input expected = TestInstance