replace-attoparsec 1.2.2.0 → 1.4.0.0
raw patch · 7 files changed
+698/−336 lines, 7 filesPVP ok
version bump matches the API change (PVP)
API changes (from Hackage documentation)
+ Replace.Attoparsec.ByteString: anyTill :: Parser a -> Parser (ByteString, a)
+ Replace.Attoparsec.ByteString: breakCap :: Parser a -> ByteString -> Maybe (ByteString, a, ByteString)
+ Replace.Attoparsec.ByteString: splitCap :: Parser a -> ByteString -> [Either ByteString a]
+ Replace.Attoparsec.Text: anyTill :: Parser a -> Parser (Text, a)
+ Replace.Attoparsec.Text: breakCap :: Parser a -> Text -> Maybe (Text, a, Text)
+ Replace.Attoparsec.Text: splitCap :: Parser a -> Text -> [Either Text a]
Files
- CHANGELOG.md +6/−0
- README.md +49/−42
- replace-attoparsec.cabal +3/−3
- src/Replace/Attoparsec/ByteString.hs +289/−132
- src/Replace/Attoparsec/Text.hs +302/−154
- tests/TestByteString.hs +24/−2
- tests/TestText.hs +25/−3
CHANGELOG.md view
@@ -1,5 +1,11 @@ # Revision history for replace-attoparsec +## 1.4.0.0 -- 2020-05-06++__Running Parsers__: Add `splitCap` and `breakCap`.++__Parser Combinators__: Add `anyTill`.+ ## 1.2.0.0 -- 2019-10-31 Benchmark improvements
README.md view
@@ -10,8 +10,8 @@ * [Benchmarks](#benchmarks) * [Hypothetically Asked Questions](#hypothetically-asked-questions) -__replace-attoparsec__ is for finding text patterns, and also editing and-replacing the found patterns.+__replace-attoparsec__ is for finding text patterns, and also+replacing or splitting on the found patterns. This activity is traditionally done with regular expressions, but __replace-attoparsec__ uses [__attoparsec__](http://hackage.haskell.org/package/attoparsec)@@ -35,6 +35,12 @@ or [`awk`](https://www.gnu.org/software/gawk/manual/gawk.html). +__replace-attoparsec__ can be used in the same sort of “string splitting”+situations in which one would use Python+[`re.split`](https://docs.python.org/3/library/re.html#re.split)+or Perl+[`split`](https://perldoc.perl.org/functions/split.html).+ See [__replace-megaparsec__](https://hackage.haskell.org/package/replace-megaparsec) for the [__megaparsec__](http://hackage.haskell.org/package/megaparsec)@@ -93,49 +99,19 @@ import Data.Char ``` -## Parsing with `sepCap` family of parser combinators--The following examples show how to match a pattern to a string of text-and separate it into sections-which match the pattern, and sections which don't match.+## Split strings with `splitCap` -### Pattern match, capture only the parsed result with `sepCap`+### Find all pattern matches, capture the matched text and the parsed result Separate the input string into sections which can be parsed as a hexadecimal number with a prefix `"0x"`, and sections which can't. Parse the numbers. ```haskell let hexparser = string "0x" *> hexadecimal :: Parser Integer-fromRight [] $ parseOnly (sepCap hexparser) "0xA 000 0xFFFF"-```-```haskell-[Right 10,Left " 000 ",Right 65535]-```--### Pattern match, capture only the matched text with `findAll`--Just get the strings sections which match the hexadecimal parser, throw away-the parsed number.--```haskell-let hexparser = string "0x" *> hexadecimal :: Parser Integer-fromRight [] $ parseOnly (findAll hexparser) "0xA 000 0xFFFF"-```-```haskell-[Right "0xA",Left " 000 ",Right "0xFFFF"]-```--### Pattern match, capture the matched text and the parsed result with `findAllCap`--Capture the parsed hexadecimal number, as well as the string section which-parses as a hexadecimal number.--```haskell-let hexparser = chunk "0x" *> hexadecimal :: Parser Integer-fromRight [] $ parseOnly (findAllCap hexparser) "0xA 000 0xFFFF"+splitCap (match hexparser) "0xA 000 0xFFFF" ``` ```haskell-[Right ("0xA",10),Left " 000 ",Right ("0xFFFF",65535)]+[Right ("0xA",10), Left " 000 ", Right ("0xFFFF",65535)] ``` ### Pattern match balanced parentheses@@ -145,7 +121,8 @@ expression. We can express the pattern with a recursive parser. ```haskell-import Data.Functor+import Data.Functor (void)+import Data.Bifunctor (second) let parens :: Parser () parens = do char '('@@ -154,13 +131,13 @@ (char ')') pure () -fromRight [] $ parseOnly (findAll parens) "(()) (()())"+second fst <$> splitCap (match parens) "(()) (()())" ``` ```haskell [Right "(())",Left " ",Right "(()())"] ``` -## Edit text strings by running parsers with `streamEdit`+## Edit text strings with `streamEdit` The following examples show how to search for a pattern in a string of text and then edit the string of text to substitute in some replacement text@@ -168,7 +145,7 @@ ### Pattern match and replace with a constant -Replace all carriage-return-newline instances with newline.+Replace all carriage-return-newline occurances with newline. ```haskell streamEdit (string "\r\n") (const "\n") "1\r\n2\r\n"@@ -204,17 +181,47 @@ "10 000 0xFFFF" ``` -### Pattern match and edit the matches with IO+### Pattern match and edit the matches with IO with `streamEditT` Find an environment variable in curly braces and replace it with its value from the environment. ```haskell-import System.Environment+import System.Environment (getEnv) streamEditT (char '{' *> manyTill anyChar (char '}')) (fmap T.pack . getEnv) "- {HOME} -" ``` ```haskell "- /home/jbrock -"+```++### Pattern match, edit the matches, and count the edits with `streamEditT`++Find and capitalize no more than three letters in a string, and return the +edited string along with the number of letters capitalized. To enable the+editor function to remember how many letters it has capitalized, we'll +run `streamEditT` in the `State` monad from the `mtl` package. Use this+technique to get the same functionality as Python+[`re.subn`](https://docs.python.org/3/library/re.html#re.subn).++```haskell+import qualified Control.Monad.State.Strict as MTL+import Control.Monad.State.Strict (get, put, runState)+import Data.Char (toUpper)++let editThree :: Char -> MTL.State Int T.Text+ editThree x = do+ i <- get+ let i' = i+1+ if i'<=3+ then do+ put i'+ pure $ T.singleton $ toUpper x+ else pure $ T.singleton x++flip runState 0 $ streamEditT (satisfy isLetter) editThree "a a a a a"+```+```haskell+("A A A a a",3) ```
replace-attoparsec.cabal view
@@ -1,7 +1,7 @@ name: replace-attoparsec-version: 1.2.2.0+version: 1.4.0.0 cabal-version: 1.18-synopsis: Find, replace, and edit text patterns with Attoparsec parsers (instead of regex)+synopsis: Find, replace, and split string patterns with Attoparsec parsers (instead of regex) homepage: https://github.com/jamesdbrock/replace-attoparsec bug-reports: https://github.com/jamesdbrock/replace-attoparsec/issues license: BSD2@@ -12,7 +12,7 @@ category: Parsing description: - Find text patterns, and also edit or replace the found patterns. Use+ Find text patterns, replace the patterns, split on the patterns. Use Attoparsec monadic parsers instead of regular expressions for pattern matching. extra-doc-files: README.md
src/Replace/Attoparsec/ByteString.hs view
@@ -4,8 +4,8 @@ -- License : BSD2 -- Maintainer: James Brock <jamesbrock@gmail.com> ----- __Replace.Attoparsec__ is for finding text patterns, and also editing and--- replacing the found patterns.+-- __Replace.Attoparsec__ is for finding text patterns, and also+-- replacing or splitting on the found patterns. -- This activity is traditionally done with regular expressions, -- but __Replace.Attoparsec__ uses "Data.Attoparsec" parsers instead for -- the pattern matching.@@ -28,6 +28,12 @@ -- or -- <https://www.gnu.org/software/gawk/manual/gawk.html awk>. --+-- __Replace.Attoparsec__ can be used in the same sort of “string splitting”+-- situations in which one would use Python+-- <https://docs.python.org/3/library/re.html#re.split re.split>+-- or Perl+-- <https://perldoc.perl.org/functions/split.html split>.+-- -- See the __[replace-attoparsec](https://hackage.haskell.org/package/replace-attoparsec)__ package README for usage examples. {-# LANGUAGE LambdaCase #-}@@ -35,14 +41,24 @@ module Replace.Attoparsec.ByteString (- -- * Parser combinator- sepCap- , findAll- , findAllCap- -- * Running parser+ --+ -- | Functions in this section are ways to run parsers. They take+ -- as arguments a @sep@ parser and some input, run the parser on the input,+ -- and return a result.+ breakCap+ , splitCap , streamEdit , streamEditT+ -- * Parser combinator+ --+ -- | Functions in this section are parser combinators. They take+ -- a @sep@ parser for an argument, combine @sep@ with another parser,+ -- and return a new parser.+ , anyTill+ , sepCap+ , findAll+ , findAllCap ) where @@ -54,15 +70,236 @@ import qualified Data.ByteString as B import qualified Data.Attoparsec.Internal.Types as AT + -- |--- == Separate and capture+-- === Break on and capture one pattern --+-- Find the first occurence of a pattern in a text stream, capture the found+-- pattern, and break the input text stream on the found pattern.+--+-- The 'breakCap' function is like 'Data.List.takeWhile', but can be predicated+-- beyond more than just the next one token. It's also like 'Data.Text.breakOn',+-- but the @needle@ can be a pattern instead of a constant string.+--+-- Be careful not to look too far+-- ahead; if the @sep@ parser looks to the end of the input then 'breakCap'+-- could be /O(n²)/.+--+-- The pattern parser @sep@ may match a zero-width pattern (a pattern which+-- consumes no parser input on success).+--+-- ==== Output+--+-- * @Nothing@ when no pattern match was found.+-- * @Just (prefix, parse_result, suffix)@ for the result of parsing the+-- pattern match, and the @prefix@ string before and the @suffix@ string+-- after the pattern match. @prefix@ and @suffix@ may be zero-length strings.+--+-- ==== Access the matched section of text+--+-- If you want to capture the matched string, then combine the pattern+-- parser @sep@ with 'Data.Attoparsec.ByteString.match'.+--+-- With the matched string, we can reconstruct in input string.+-- For all @input@, @sep@, if+--+-- @+-- let ('Just' (prefix, (infix, _), suffix)) = breakCap ('Data.Attoparsec.ByteString.match' sep) input+-- @+--+-- then+--+-- @+-- input == prefix '<>' infix '<>' suffix+-- @+breakCap+ :: Parser a+ -- ^ The pattern matching parser @sep@+ -> B.ByteString+ -- ^ The input stream of text+ -> Maybe (B.ByteString, a, B.ByteString)+ -- ^ Maybe (prefix, parse_result, suffix)+breakCap sep input =+ case parseOnly pser input of+ (Left _) -> Nothing+ (Right x) -> Just x+ where+ pser = do+ (prefix, cap) <- anyTill sep+ suffix <- A.takeByteString+ pure (prefix, cap, suffix)+{-# INLINABLE breakCap #-}+++-- |+-- === Split on and capture all patterns+--+-- Find all occurences of the pattern @sep@, split the input string, capture+-- all the patterns and the splits.+--+-- The input string will be split on every leftmost non-overlapping occurence+-- of the pattern @sep@. The output list will contain+-- the parsed result of input string sections which match the @sep@ pattern+-- in 'Right', and non-matching sections in 'Left'.+--+-- 'splitCap' depends on 'sepCap', see 'sepCap' for more details.+--+-- ==== Access the matched section of text+--+-- If you want to capture the matched strings, then combine the pattern+-- parser @sep@ with 'Data.Attoparsec.ByteString.match'.+--+-- With the matched strings, we can reconstruct in input string.+-- For all @input@, @sep@, if+--+-- @+-- let output = splitCap ('Data.Attoparsec.ByteString.match' sep) input+-- @+--+-- then+--+-- @+-- input == 'Data.Monoid.mconcat' ('Data.Bifunctor.second' 'Data.Tuple.fst' '<$>' output)+-- @+splitCap+ :: Parser a+ -- ^ The pattern matching parser @sep@+ -> B.ByteString+ -- ^ The input stream of text+ -> [Either B.ByteString a]+ -- ^ List of matching and non-matching input sections+splitCap sep input = do+ case parseOnly (sepCap sep) input of+ (Left _) -> undefined -- sepCap can never fail+ (Right r) -> r+{-# INLINABLE splitCap #-}+++-- |+-- === Stream editor+--+-- Also known as “find-and-replace”, or “match-and-substitute”. Finds all+-- of the sections of the stream which match the pattern @sep@, and replaces+-- them with the result of the @editor@ function.+--+-- ==== Access the matched section of text in the @editor@+--+-- If you want access to the matched string in the @editor@ function,+-- then combine the pattern parser @sep@+-- with 'Data.Attoparsec.ByteString.match'. This will effectively change+-- the type of the @editor@ function to @(ByteString,a) -> ByteString@.+--+-- This allows us to write an @editor@ function which can choose to not+-- edit the match and just leave it as it is. If the @editor@ function+-- returns the first item in the tuple, then @streamEdit@ will not change+-- the matched string.+--+-- So, for all @sep@:+--+-- @+-- streamEdit ('Data.Attoparsec.ByteString.match' sep) 'Data.Tuple.fst' ≡ 'Data.Function.id'+-- @+streamEdit+ :: Parser a+ -- ^ The pattern matching parser @sep@+ -> (a -> B.ByteString)+ -- ^ The @editor@ function. Takes a parsed result of @sep@+ -- and returns a new stream section for the replacement.+ -> B.ByteString+ -- ^ The input stream of text to be edited+ -> B.ByteString+ -- ^ The edited input stream+streamEdit sep editor = runIdentity . streamEditT sep (Identity . editor)+{-# INLINABLE streamEdit #-}+++-- |+-- === Stream editor transformer+--+-- Monad transformer version of 'streamEdit'.+--+-- The @editor@ function will run in the underlying monad context.+--+-- If you want to do 'IO' operations in the @editor@ function then+-- run this in 'IO'.+--+-- If you want the @editor@ function to remember some state,+-- then run this in a stateful monad.+streamEditT+ :: (Monad m)+ => Parser a+ -- ^ The pattern matching parser @sep@+ -> (a -> m B.ByteString)+ -- ^ The @editor@ function. Takes a parsed result of @sep@+ -- and returns a new stream section for the replacement.+ -> B.ByteString+ -- ^ The input stream of text to be edited+ -> m B.ByteString+ -- ^ The edited input stream+streamEditT sep editor input = do+ case parseOnly (sepCap sep) input of+ (Left err) -> error err+ -- this function should never error, because it only errors+ -- when the 'sepCap' parser fails, and the 'sepCap' parser+ -- can never fail. If this function ever throws an error, please+ -- report that as a bug.+ -- (We don't use MonadFail because Identity is not a MonadFail.)+ (Right r) -> fmap mconcat $ traverse (either return editor) r+{-# INLINABLE streamEditT #-}++++-- |+-- === Specialized <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html#v:manyTill_ manyTill_>+--+-- Parser combinator to consume and capture input until the @sep@ pattern+-- matches, equivalent to+-- @'Control.Monad.Combinators.manyTill_' 'Data.Attoparsec.ByteString.anyWord8' sep@.+-- On success, returns the prefix before the pattern match and the parsed match.+--+-- @sep@ may be a zero-width parser, it may succeed without consuming any+-- input.+--+-- This combinator will produce a parser which acts+-- like 'Data.Attoparsec.ByteString.takeTill' but is predicated beyond more than+-- just the next one token. It is also like+-- 'Data.Attoparsec.ByteString.takeTill' in that it is a “high performance”+-- parser.+anyTill+ :: Parser a -- ^ The pattern matching parser @sep@+ -> Parser (B.ByteString, a) -- ^ parser+anyTill sep = do+ begin <- getOffset+ (end, x) <- go+ prefix <- substring begin end+ pure (prefix, x)+ where+ go = do+ end <- getOffset+ r <- optional sep+ case r of+ Nothing -> atEnd >>= \case+ True -> empty+ False -> advance >> go+ Just x -> pure (end, x)+++-- |+-- === Separate and capture+-- -- Parser combinator to find all of the non-overlapping ocurrences -- of the pattern @sep@ in a text stream. -- The 'sepCap' parser will always consume its entire input and can never fail. ----- === Output+-- @sepCap@ is similar to the @sep*@ family of functions found in+-- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html parser-combinators>+-- and+-- <http://hackage.haskell.org/package/parsers/docs/Text-Parser-Combinators.html parsers>,+-- but it returns the parsed result of the @sep@ parser instead+-- of throwing it away. --+-- ==== Output+-- -- The input stream is separated and output into a list of sections: -- -- * Sections which can parsed by the pattern @sep@ will be parsed and captured@@ -76,7 +313,7 @@ -- the entire input stream will be returned as one non-matching 'Left' section. -- * The output list will not contain two consecutive 'Left' sections. ----- === Zero-width matches forbidden+-- ==== Zero-width matches forbidden -- -- If the pattern matching parser @sep@ would succeed without consuming any -- input then 'sepCap' will force it to fail.@@ -84,22 +321,9 @@ -- then it can match the same zero-width pattern again at the same position -- on the next iteration, which would result in an infinite number of -- overlapping pattern matches.------ === Notes------ This @sepCap@ parser combinator is the basis for all of the other--- features of this module.------ It is similar to the @sep*@ family of functions found in--- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html parser-combinators>--- and--- <http://hackage.haskell.org/package/parsers/docs/Text-Parser-Combinators.html parsers>--- but, importantly, it returns the parsed result of the @sep@ parser instead--- of throwing it away, like--- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html#v:manyTill_ manyTill_>. sepCap :: Parser a -- ^ The pattern matching parser @sep@- -> Parser [Either B.ByteString a]+ -> Parser [Either B.ByteString a] -- ^ parser sepCap sep = getOffset >>= go where -- the go function will search for the first pattern match,@@ -144,41 +368,11 @@ (Right x:) <$> go offsetAfter Nothing -> go offsetBegin -- no match, try again )- -- Using this advance function instead of 'anyWord8' seems to give us- -- a 5%-20% performance improvement.- --- -- It's safe to use 'advance' because after 'advance' we always check- -- for 'endOfInput' before trying to read anything from the buffer.- --- -- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.ByteString.Internal.html#anyWord8- -- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.ByteString.Internal.html#advance- -- advance :: Parser ()- advance = AT.Parser $ \t pos more _lose succes ->- succes t (pos + AT.Pos 1) more ()-- -- Extract a substring from part of the buffer that we've already visited.- -- Does not check bounds.- --- -- The idea here is that we go back and run the parser 'take' at the Pos- -- which we saved from before, and then we continue from the current Pos,- -- hopefully without messing up the internal parser state.- --- -- Should be equivalent to the unexported function- -- Data.Attoparsec.ByteString.Buffer.substring- -- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.ByteString.Buffer.html#substring- --- -- This is a performance optimization for gathering the unmatched- -- sections of the input. The alternative is to accumulate unmatched- -- characters one anyWord8 at a time in a list of [Word8] and then pack- -- them into a ByteString.- substring :: Int -> Int -> Parser B.ByteString- substring !pos1 !pos2 = AT.Parser $ \t pos more lose succes ->- let succes' _t _pos _more a = succes t pos more a- in AT.runParser (A.take (pos2 - pos1)) t (AT.Pos pos1) more lose succes' {-# INLINABLE sepCap #-} + -- |--- == Find all occurences, parse and capture pattern matches+-- === Find all occurences, parse and capture pattern matches -- -- Parser combinator for finding all occurences of a pattern in a stream. --@@ -193,13 +387,13 @@ -- @ findAllCap :: Parser a -- ^ The pattern matching parser @sep@- -> Parser [Either B.ByteString (B.ByteString, a)]+ -> Parser [Either B.ByteString (B.ByteString, a)] -- ^ parser findAllCap sep = sepCap (match sep) {-# INLINABLE findAllCap #-} -- |--- == Find all occurences+-- === Find all occurences -- -- Parser combinator for finding all occurences of a pattern in a stream. --@@ -214,90 +408,53 @@ -- @ findAll :: Parser a -- ^ The pattern matching parser @sep@- -> Parser [Either B.ByteString B.ByteString]+ -> Parser [Either B.ByteString B.ByteString] -- ^ parser findAll sep = (fmap.fmap) (second fst) $ sepCap (match sep) {-# INLINABLE findAll #-} -- |--- == Stream editor------ Also known as “find-and-replace”, or “match-and-substitute”. Finds all--- of the sections of the stream which match the pattern @sep@, and replaces--- them with the result of the @editor@ function.------ This function is not a “parser combinator,” it is--- a “way to run a parser”, like 'Data.Attoparsec.ByteString.parse'--- or 'Data.Attoparsec.ByteString.parseOnly'.------ === Access the matched section of text in the @editor@------ If you want access to the matched string in the @editor@ function,--- then combine the pattern parser @sep@--- with 'Data.Attoparsec.ByteString.match'. This will effectively change--- the type of the @editor@ function to @(ByteString,a) -> ByteString@.------ This allows us to write an @editor@ function which can choose to not--- edit the match and just leave it as it is. If the @editor@ function--- returns the first item in the tuple, then @streamEdit@ will not change--- the matched string.------ So, for all @sep@:+-- Get the 'Data.Attoparsec.Internal.Types.Parser' current offset+-- 'Data.Attoparsec.Internal.Types.Pos' in the stream. ----- @--- streamEdit ('Data.Attoparsec.ByteString.match' sep) 'Data.Tuple.fst' ≡ 'Data.Function.id'--- @-streamEdit- -- :: forall s a. (Stream s, Monoid s, Tokens s ~ s, Show s, Show (Token s), Typeable s)- :: Parser a- -- ^ The parser @sep@ for the pattern of interest.- -> (a -> B.ByteString)- -- ^ The @editor@ function. Takes a parsed result of @sep@- -- and returns a new stream section for the replacement.- -> B.ByteString- -- ^ The input stream of text to be edited.- -> B.ByteString-streamEdit sep editor = runIdentity . streamEditT sep (Identity . editor)-{-# INLINABLE streamEdit #-}+-- [“… you know you're in an uncomfortable state of sin :-)” — bos](https://github.com/bos/attoparsec/issues/101)+getOffset :: Parser Int+getOffset = AT.Parser $ \t pos more _ succ' -> succ' t pos more (AT.fromPos pos) + -- |--- == Stream editor transformer------ Monad transformer version of 'streamEdit'.------ The @editor@ function will run in the underlying monad context.+-- Using this advance function instead of 'anyWord8' seems to give us+-- a 5%-20% performance improvement for sepCap. ----- If you want to do 'IO' operations in the @editor@ function then--- run this in 'IO'.+-- It's safe to use 'advance' because after 'advance' we always check+-- for 'endOfInput' before trying to read anything from the buffer. ----- If you want the @editor@ function to remember some state,--- then run this in a stateful monad.-streamEditT- :: (Monad m)- => Parser a- -- ^ The parser @sep@ for the pattern of interest.- -> (a -> m B.ByteString)- -- ^ The @editor@ function. Takes a parsed result of @sep@- -- and returns a new stream section for the replacement.- -> B.ByteString- -- ^ The input stream of text to be edited.- -> m B.ByteString-streamEditT sep editor input = do- case parseOnly (sepCap sep) input of- (Left err) -> error err- -- this function should never error, because it only errors- -- when the 'sepCap' parser fails, and the 'sepCap' parser- -- can never fail. If this function ever throws an error, please- -- report that as a bug.- -- (We don't use MonadFail because Identity is not a MonadFail.)- (Right r) -> fmap mconcat $ traverse (either return editor) r-{-# INLINABLE streamEditT #-}+-- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.ByteString.Internal.html#anyWord8+-- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.ByteString.Internal.html#advance+advance :: Parser ()+advance = AT.Parser $ \t pos more _lose succes ->+ succes t (pos + AT.Pos 1) more ()+{-# INLINABLE advance #-} + -- |--- Get the 'Data.Attoparsec.Internal.Types.Parser' current offset--- 'Data.Attoparsec.Internal.Types.Pos' in the stream.+-- Extract a substring from part of the buffer that we've already visited.+-- Does not check bounds. ----- [“… you know you're in an uncomfortable state of sin :-)” — bos](https://github.com/bos/attoparsec/issues/101)-getOffset :: Parser Int-getOffset = AT.Parser $ \t pos more _ succ' -> succ' t pos more (AT.fromPos pos)-+-- The idea here is that we go back and run the parser 'take' at the Pos+-- which we saved from before, and then we continue from the current Pos,+-- hopefully without messing up the internal parser state.+--+-- Should be equivalent to the unexported function+-- Data.Attoparsec.ByteString.Buffer.substring+-- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.ByteString.Buffer.html#substring+--+-- This is a performance optimization for gathering the unmatched+-- sections of the input. The alternative is to accumulate unmatched+-- characters one anyWord8 at a time in a list of [Word8] and then pack+-- them into a ByteString.+substring :: Int -> Int -> Parser B.ByteString+substring !pos1 !pos2 = AT.Parser $ \t pos more lose succes ->+ let succes' _t _pos _more a = succes t pos more a+ in AT.runParser (A.take (pos2 - pos1)) t (AT.Pos pos1) more lose succes'+{-# INLINABLE substring #-}
src/Replace/Attoparsec/Text.hs view
@@ -4,8 +4,8 @@ -- License : BSD2 -- Maintainer: James Brock <jamesbrock@gmail.com> ----- __Replace.Attoparsec__ is for finding text patterns, and also editing and--- replacing the found patterns.+-- __Replace.Attoparsec__ is for finding text patterns, and also+-- replacing or splitting on the found patterns. -- This activity is traditionally done with regular expressions, -- but __Replace.Attoparsec__ uses "Data.Attoparsec" parsers instead for -- the pattern matching.@@ -28,6 +28,12 @@ -- or -- <https://www.gnu.org/software/gawk/manual/gawk.html awk>. --+-- __Replace.Attoparsec__ can be used in the same sort of “string splitting”+-- situations in which one would use Python+-- <https://docs.python.org/3/library/re.html#re.split re.split>+-- or Perl+-- <https://perldoc.perl.org/functions/split.html split>.+-- -- See the __[replace-attoparsec](https://hackage.haskell.org/package/replace-attoparsec)__ package README for usage examples. {-# LANGUAGE LambdaCase #-}@@ -35,14 +41,24 @@ module Replace.Attoparsec.Text (- -- * Parser combinator- sepCap- , findAll- , findAllCap- -- * Running parser+ --+ -- | Functions in this section are ways to run parsers. They take+ -- as arguments a @sep@ parser and some input, run the parser on the input,+ -- and return a result.+ breakCap+ , splitCap , streamEdit , streamEditT+ -- * Parser combinator+ --+ -- | Functions in this section are parser combinators. They take+ -- a @sep@ parser for an argument, combine @sep@ with another parser,+ -- and return a new parser.+ , anyTill+ , sepCap+ , findAll+ , findAllCap ) where @@ -55,15 +71,234 @@ import qualified Data.Text.Internal as TI import qualified Data.Attoparsec.Internal.Types as AT + -- |--- == Separate and capture+-- === Break on and capture one pattern --+-- Find the first occurence of a pattern in a text stream, capture the found+-- pattern, and break the input text stream on the found pattern.+--+-- The 'breakCap' function is like 'Data.List.takeWhile', but can be predicated+-- beyond more than just the next one token. It's also like 'Data.Text.breakOn',+-- but the @needle@ can be a pattern instead of a constant string.+--+-- Be careful not to look too far+-- ahead; if the @sep@ parser looks to the end of the input then 'breakCap'+-- could be /O(n²)/.+--+-- The pattern parser @sep@ may match a zero-width pattern (a pattern which+-- consumes no parser input on success).+--+-- ==== Output+--+-- * @Nothing@ when no pattern match was found.+-- * @Just (prefix, parse_result, suffix)@ for the result of parsing the+-- pattern match, and the @prefix@ string before and the @suffix@ string+-- after the pattern match. @prefix@ and @suffix@ may be zero-length strings.+--+-- ==== Access the matched section of text+--+-- If you want to capture the matched string, then combine the pattern+-- parser @sep@ with 'Data.Attoparsec.Text.match'.+--+-- With the matched string, we can reconstruct in input string.+-- For all @input@, @sep@, if+--+-- @+-- let ('Just' (prefix, (infix, _), suffix)) = breakCap ('Data.Attoparsec.Text.match' sep) input+-- @+--+-- then+--+-- @+-- input == prefix '<>' infix '<>' suffix+-- @+breakCap+ :: Parser a+ -- ^ The pattern matching parser @sep@+ -> T.Text+ -- ^ The input stream of text+ -> Maybe (T.Text, a, T.Text)+ -- ^ Maybe (prefix, parse_result, suffix)+breakCap sep input =+ case parseOnly pser input of+ (Left _) -> Nothing+ (Right x) -> Just x+ where+ pser = do+ (prefix, cap) <- anyTill sep+ suffix <- A.takeText+ pure (prefix, cap, suffix)+{-# INLINABLE breakCap #-}++--+-- |+-- === Split on and capture all patterns+--+-- Find all occurences of the pattern @sep@, split the input string, capture+-- all the patterns and the splits.+--+-- The input string will be split on every leftmost non-overlapping occurence+-- of the pattern @sep@. The output list will contain+-- the parsed result of input string sections which match the @sep@ pattern+-- in 'Right', and non-matching sections in 'Left'.+--+-- 'splitCap' depends on 'sepCap', see 'sepCap' for more details.+--+-- ==== Access the matched section of text+--+-- If you want to capture the matched strings, then combine the pattern+-- parser @sep@ with 'Data.Attoparsec.Text.match'.+--+-- With the matched strings, we can reconstruct in input string.+-- For all @input@, @sep@, if+--+-- @+-- let output = splitCap ('Data.Attoparsec.Text.match' sep) input+-- @+--+-- then+--+-- @+-- input == 'Data.Monoid.mconcat' ('Data.Bifunctor.second' 'Data.Tuple.fst' '<$>' output)+-- @+splitCap+ :: Parser a+ -- ^ The pattern matching parser @sep@+ -> T.Text+ -- ^ The input stream of text+ -> [Either T.Text a]+ -- ^ List of matching and non-matching input sections+splitCap sep input = do+ case parseOnly (sepCap sep) input of+ (Left _) -> undefined -- sepCap can never fail+ (Right r) -> r+{-# INLINABLE splitCap #-}+++-- |+-- === Stream editor+--+-- Also known as “find-and-replace”, or “match-and-substitute”. Finds all+-- of the sections of the stream which match the pattern @sep@, and replaces+-- them with the result of the @editor@ function.+--+-- ==== Access the matched section of text in the @editor@+--+-- If you want access to the matched string in the @editor@ function,+-- then combine the pattern parser @sep@+-- with 'Data.Attoparsec.Text.match'. This will effectively change+-- the type of the @editor@ function to @(Text,a) -> Text@.+--+-- This allows us to write an @editor@ function which can choose to not+-- edit the match and just leave it as it is. If the @editor@ function+-- returns the first item in the tuple, then @streamEdit@ will not change+-- the matched string.+--+-- So, for all @sep@:+--+-- @+-- streamEdit ('Data.Attoparsec.Text.match' sep) 'Data.Tuple.fst' ≡ 'Data.Function.id'+-- @+streamEdit+ :: Parser a+ -- ^ The pattern matching parser @sep@+ -> (a -> T.Text)+ -- ^ The @editor@ function. Takes a parsed result of @sep@+ -- and returns a new stream section for the replacement.+ -> T.Text+ -- ^ The input stream of text to be edited+ -> T.Text+ -- ^ The edited input stream+streamEdit sep editor = runIdentity . streamEditT sep (Identity . editor)+{-# INLINABLE streamEdit #-}+++-- |+-- === Stream editor transformer+--+-- Monad transformer version of 'streamEdit'.+--+-- The @editor@ function will run in the underlying monad context.+--+-- If you want to do 'IO' operations in the @editor@ function then+-- run this in 'IO'.+--+-- If you want the @editor@ function to remember some state,+-- then run this in a stateful monad.+streamEditT+ :: (Monad m)+ => Parser a+ -- ^ The pattern matching parser @sep@+ -> (a -> m T.Text)+ -- ^ The @editor@ function. Takes a parsed result of @sep@+ -- and returns a new stream section for the replacement.+ -> T.Text+ -- ^ The input stream of text to be edited+ -> m T.Text+ -- ^ The edited input stream+streamEditT sep editor input = do+ case parseOnly (sepCap sep) input of+ (Left err) -> error err+ -- this function should never error, because it only errors+ -- when the 'sepCap' parser fails, and the 'sepCap' parser+ -- can never fail. If this function ever throws an error, please+ -- report that as a bug.+ -- (We don't use MonadFail because Identity is not a MonadFail.)+ (Right r) -> fmap mconcat $ traverse (either return editor) r+{-# INLINABLE streamEditT #-}+++-- |+-- === Specialized <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html#v:manyTill_ manyTill_>+--+-- Parser combinator to consume and capture input until the @sep@ pattern+-- matches, equivalent to+-- @'Control.Monad.Combinators.manyTill_' 'Data.Attoparsec.Text.anyChar' sep@.+-- On success, returns the prefix before the pattern match and the parsed match.+--+-- @sep@ may be a zero-width parser, it may succeed without consuming any+-- input.+--+-- This combinator will produce a parser which acts+-- like 'Data.Attoparsec.Text.takeTill' but is predicated beyond more than+-- just the next one token. It is also like+-- 'Data.Attoparsec.Text.takeTill' in that it is a “high performance” parser.+anyTill+ :: Parser a -- ^ The pattern matching parser @sep@+ -> Parser (T.Text, a) -- ^ parser+anyTill sep = do+ begin <- getOffset+ (end, x) <- go+ prefix <- substring begin end+ pure (prefix, x)+ where+ go = do+ end <- getOffset+ r <- optional sep+ case r of+ Nothing -> atEnd >>= \case+ True -> empty+ False -> anyChar >> go+ Just x -> pure (end, x)+++-- |+-- === Separate and capture+-- -- Parser combinator to find all of the non-overlapping ocurrences -- of the pattern @sep@ in a text stream. -- The 'sepCap' parser will always consume its entire input and can never fail. ----- === Output+-- 'sepCap' is similar to the @sep*@ family of functions found in+-- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html parser-combinators>+-- and+-- <http://hackage.haskell.org/package/parsers/docs/Text-Parser-Combinators.html parsers>,+-- but it returns the parsed result of the @sep@ parser instead+-- of throwing it away. --+-- ==== Output+-- -- The input stream is separated and output into a list of sections: -- -- * Sections which can parsed by the pattern @sep@ will be parsed and captured@@ -77,7 +312,7 @@ -- the entire input stream will be returned as one non-matching 'Left' section. -- * The output list will not contain two consecutive 'Left' sections. ----- === Zero-width matches forbidden+-- ==== Zero-width matches forbidden -- -- If the pattern matching parser @sep@ would succeed without consuming any -- input then 'sepCap' will force it to fail.@@ -85,24 +320,9 @@ -- then it can match the same zero-width pattern again at the same position -- on the next iteration, which would result in an infinite number of -- overlapping pattern matches.------ === Notes------ This @sepCap@ parser combinator is the basis for all of the other--- features of this module.------ It is similar to the @sep*@ family of functions found in--- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html parser-combinators>--- and--- <http://hackage.haskell.org/package/parsers/docs/Text-Parser-Combinators.html parsers>--- but, importantly, it returns the parsed result of the @sep@ parser instead--- of throwing it away, like--- <http://hackage.haskell.org/package/parser-combinators/docs/Control-Monad-Combinators.html#v:manyTill_ manyTill_>.---- sepCap :: Parser a -- ^ The pattern matching parser @sep@- -> Parser [Either T.Text a]+ -> Parser [Either T.Text a] -- ^ parser sepCap sep = getOffset >>= go where -- the go function will search for the first pattern match,@@ -147,61 +367,11 @@ (Right x:) <$> go offsetAfter Nothing -> go offsetBegin -- no match, try again )-- -- Extract a substring from part of the buffer that we've already visited.- --- -- The idea here is that we go back and run the parser 'take' at the Pos- -- which we saved from before, and then we continue from the current Pos,- -- hopefully without messing up the internal parser state.- -- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.Text.Internal.html#take- --- -- Should be equivalent to the unexported function- -- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.Text.Internal.html#substring- --- -- This is a performance optimization for gathering the unmatched- -- sections of the input. The alternative is to accumulate unmatched- -- characters one anyChar at a time in a list of [Char] and then pack- -- them into a Text.- substring :: Int -> Int -> Parser T.Text- substring !bgn !end = AT.Parser $ \t pos more lose succes ->- let succes' _t _pos _more a = succes t pos more a- in- AT.runParser (takeCheat (end - bgn)) t (AT.Pos bgn) more lose succes'- where- -- Dear reader, you deserve an explanation for 'takeCheat'. The- -- alternative to running 'takeCheat' here would be the following line:- --- -- AT.runParser (A.take (end - bgn)) t (AT.Pos bgn) more lose succes'- --- -- But 'Attoparsec.take' is not correct, and 'takeCheat' is correct.- -- It is correct because the Pos which we got from 'getOffset' is an- -- index into the underlying Data.Text.Array, so (end - bgn) is- -- in units of the length of the Data.Text.Array, not in units of the- -- number of Chars.- --- -- Furthermore 'takeCheat' is a lot faster because 'A.take' takes a- -- number of Chars and then iterates over the Text by the number- -- of Chars, advancing by 4 bytes when it encounters a wide Char.- -- So, O(N). takeCheat is O(1).- --- -- This will be fine as long as we always call 'takeCheat' on the- -- immutable, already-visited part of the Attoparsec.Text.Buffer's- -- Data.Text.Array. Which we do.- --- -- It's named 'takeCheat' because we're getting access to- -- the Attoparsec.Text.Buffer through the Data.Text.Internal- -- interface, even though Attoparsec is extremely vigilant about- -- not exposing its buffers.- --- -- http://hackage.haskell.org/package/text-1.2.3.1/docs/Data-Text-Internal.html- -- takeCheat :: Int -> Parser T.Text- takeCheat len = do- (TI.Text arr off _len) <- A.take 1- return (TI.Text arr off len) {-# INLINABLE sepCap #-} + -- |--- == Find all occurences, parse and capture pattern matches+-- === Find all occurences, parse and capture pattern matches -- -- Parser combinator for finding all occurences of a pattern in a stream. --@@ -216,12 +386,12 @@ -- @ findAllCap :: Parser a -- ^ The pattern matching parser @sep@- -> Parser [Either T.Text (T.Text, a)]+ -> Parser [Either T.Text (T.Text, a)] -- ^ parser findAllCap sep = sepCap (match sep) {-# INLINABLE findAllCap #-} -- |--- == Find all occurences+-- === Find all occurences -- -- Parser combinator for finding all occurences of a pattern in a stream. --@@ -236,84 +406,11 @@ -- @ findAll :: Parser a -- ^ The pattern matching parser @sep@- -> Parser [Either T.Text T.Text]+ -> Parser [Either T.Text T.Text] -- ^ parser findAll sep = (fmap.fmap) (second fst) $ sepCap (match sep) {-# INLINABLE findAll #-} --- |--- == Stream editor------ Also known as “find-and-replace”, or “match-and-substitute”. Finds all--- of the sections of the stream which match the pattern @sep@, and replaces--- them with the result of the @editor@ function.------ This function is not a “parser combinator,” it is--- a “way to run a parser”, like 'Data.Attoparsec.Text.parse'--- or 'Data.Attoparsec.Text.parseOnly'.------ === Access the matched section of text in the @editor@------ If you want access to the matched string in the @editor@ function,--- then combine the pattern parser @sep@--- with 'Data.Attoparsec.Text.match'. This will effectively change--- the type of the @editor@ function to @(Text,a) -> Text@.------ This allows us to write an @editor@ function which can choose to not--- edit the match and just leave it as it is. If the @editor@ function--- returns the first item in the tuple, then @streamEdit@ will not change--- the matched string.------ So, for all @sep@:------ @--- streamEdit ('Data.Attoparsec.Text.match' sep) 'Data.Tuple.fst' ≡ 'Data.Function.id'--- @-streamEdit- :: Parser a- -- ^ The parser @sep@ for the pattern of interest.- -> (a -> T.Text)- -- ^ The @editor@ function. Takes a parsed result of @sep@- -- and returns a new stream section for the replacement.- -> T.Text- -- ^ The input stream of text to be edited.- -> T.Text-streamEdit sep editor = runIdentity . streamEditT sep (Identity . editor)-{-# INLINABLE streamEdit #-} --- |--- == Stream editor transformer------ Monad transformer version of 'streamEdit'.------ The @editor@ function will run in the underlying monad context.------ If you want to do 'IO' operations in the @editor@ function then--- run this in 'IO'.------ If you want the @editor@ function to remember some state,--- then run this in a stateful monad.-streamEditT- :: (Monad m)- => Parser a- -- ^ The parser @sep@ for the pattern of interest.- -> (a -> m T.Text)- -- ^ The @editor@ function. Takes a parsed result of @sep@- -- and returns a new stream section for the replacement.- -> T.Text- -- ^ The input stream of text to be edited.- -> m T.Text-streamEditT sep editor input = do- case parseOnly (sepCap sep) input of- (Left err) -> error err- -- this function should never error, because it only errors- -- when the 'sepCap' parser fails, and the 'sepCap' parser- -- can never fail. If this function ever throws an error, please- -- report that as a bug.- -- (We don't use MonadFail because Identity is not a MonadFail.)- (Right r) -> fmap mconcat $ traverse (either return editor) r-{-# INLINABLE streamEditT #-}---- | -- Get the 'Data.Attoparsec.Internal.Types.Parser' current offset -- 'Data.Attoparsec.Internal.Types.Pos' in the stream. --@@ -327,3 +424,54 @@ getOffset = AT.Parser $ \t pos more _ succ' -> succ' t pos more (AT.fromPos pos) {-# INLINABLE getOffset #-} ++-- Extract a substring from part of the buffer that we've already visited.+--+-- The idea here is that we go back and run the parser 'take' at the Pos+-- which we saved from before, and then we continue from the current Pos,+-- hopefully without messing up the internal parser state.+-- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.Text.Internal.html#take+--+-- Should be equivalent to the unexported function+-- http://hackage.haskell.org/package/attoparsec-0.13.2.3/docs/src/Data.Attoparsec.Text.Internal.html#substring+--+-- This is a performance optimization for gathering the unmatched+-- sections of the input. The alternative is to accumulate unmatched+-- characters one anyChar at a time in a list of [Char] and then pack+-- them into a Text.+substring :: Int -> Int -> Parser T.Text+substring !bgn !end = AT.Parser $ \t pos more lose succes ->+ let succes' _t _pos _more a = succes t pos more a+ in+ AT.runParser (takeCheat (end - bgn)) t (AT.Pos bgn) more lose succes'+ where+ -- Dear reader, you deserve an explanation for 'takeCheat'. The+ -- alternative to running 'takeCheat' here would be the following line:+ --+ -- AT.runParser (A.take (end - bgn)) t (AT.Pos bgn) more lose succes'+ --+ -- But 'Attoparsec.take' is not correct, and 'takeCheat' is correct.+ -- It is correct because the Pos which we got from 'getOffset' is an+ -- index into the underlying Data.Text.Array, so (end - bgn) is+ -- in units of the length of the Data.Text.Array, not in units of the+ -- number of Chars.+ --+ -- Furthermore 'takeCheat' is a lot faster because 'A.take' takes a+ -- number of Chars and then iterates over the Text by the number+ -- of Chars, advancing by 4 bytes when it encounters a wide Char.+ -- So, O(N). takeCheat is O(1).+ --+ -- This will be fine as long as we always call 'takeCheat' on the+ -- immutable, already-visited part of the Attoparsec.Text.Buffer's+ -- Data.Text.Array. Which we do.+ --+ -- It's named 'takeCheat' because we're getting access to+ -- the Attoparsec.Text.Buffer through the Data.Text.Internal+ -- interface, even though Attoparsec is extremely vigilant about+ -- not exposing its buffers.+ --+ -- http://hackage.haskell.org/package/text-1.2.3.1/docs/Data-Text-Internal.html+ -- takeCheat :: Int -> Parser T.Text+ takeCheat len = do+ (TI.Text arr off _len) <- A.take 0+ return (TI.Text arr off len)
tests/TestByteString.hs view
@@ -7,6 +7,7 @@ import Distribution.TestSuite as TestSuite import Data.Attoparsec.ByteString as A+import Data.Attoparsec.Combinator import qualified Data.ByteString as B import Data.ByteString.Internal (c2w) import "parsers" Text.Parser.Token@@ -48,6 +49,14 @@ , Test $ streamEditTest "x to o inner" (string "x") (const "o") " x x x " " o o o " , Test $ streamEditTest "ordering" (string "456") (const "ABC") "123456789" "123ABC789" , Test $ streamEditTest "empty input" (match (fail "")) (fst) "" ""+ , Test $ breakCapTest "basic" upperChar "aAa" (Just ("a", c2w 'A', "a"))+ , Test $ breakCapTest "first" upperChar "Aa" (Just ("", c2w 'A', "a"))+ , Test $ breakCapTest "last" upperChar "aA" (Just ("a", c2w 'A', ""))+ , Test $ breakCapTest "fail" upperChar "aaa" Nothing+ , Test $ breakCapTest "match" (match upperChar) "aAa" (Just ("a", ("A",c2w 'A'), "a"))+ , Test $ breakCapTest "zero-width" (lookAhead upperChar) "aAa" (Just ("a", c2w 'A', "Aa"))+ , Test $ breakCapTest "empty input" upperChar "" Nothing+ , Test $ breakCapTest "empty input zero-width" (return () :: Parser ()) "" (Just ("", (), "")) ] where@@ -59,7 +68,7 @@ if (output == expected) then return (Finished Pass) else return (Finished $ TestSuite.Fail- $ show output ++ " ≠ " ++ show expected)+ $ "got " <> show output <> " expected " <> show expected) , name = "parseOnly sepCap " <> nam , tags = [] , options = []@@ -79,7 +88,7 @@ if (output == expected) then return (Finished Pass) else return (Finished $ TestSuite.Fail- $ show output ++ " ≠ " ++ show expected)+ $ "got " <> show output <> " expected " <> show expected) A.Done _i _output -> return (Finished $ TestSuite.Fail $ "Should ask for more input") A.Done _i _output -> return (Finished $ TestSuite.Fail $ "Should ask for more input") , name = "parse Partial sepCap " <> nam@@ -96,6 +105,19 @@ else return (Finished $ TestSuite.Fail $ show output ++ " ≠ " ++ show expected) , name = "streamEdit " ++ nam+ , tags = []+ , options = []+ , setOption = \_ _ -> Left "no options supported"+ }++ breakCapTest nam sep input expected = TestInstance+ { run = do+ let output = breakCap sep input+ if (output == expected)+ then return (Finished Pass)+ else return (Finished $ TestSuite.Fail+ $ "got " <> show output <> " expected " <> show expected)+ , name = "breakCap " ++ nam , tags = [] , options = [] , setOption = \_ _ -> Left "no options supported"
tests/TestText.hs view
@@ -7,6 +7,7 @@ import Distribution.TestSuite as TestSuite import Replace.Attoparsec.Text import Data.Attoparsec.Text as A+import Data.Attoparsec.Combinator import qualified Data.Text as T import Text.Parser.Char (upper) import Control.Applicative@@ -54,6 +55,14 @@ , Test $ streamEditTest "x to o inner" (string "x") (const "o") " x x x " " o o o " , Test $ streamEditTest "ordering" (string "456") (const "ABC") "123456789" "123ABC789" , Test $ streamEditTest "empty input" (match (fail "")) (fst) "" ""+ , Test $ breakCapTest "basic" upper "aAa" (Just ("a", 'A', "a"))+ , Test $ breakCapTest "first" upper "Aa" (Just ("", 'A', "a"))+ , Test $ breakCapTest "last" upper "aA" (Just ("a", 'A', ""))+ , Test $ breakCapTest "fail" upper "aaa" Nothing+ , Test $ breakCapTest "match" (match upper) "aAa" (Just ("a", ("A",'A'), "a"))+ , Test $ breakCapTest "zero-width" (lookAhead upper) "aAa" (Just ("a",'A', "Aa"))+ , Test $ breakCapTest "empty input" upper "" Nothing+ , Test $ breakCapTest "empty input zero-width" (return () :: Parser ()) "" (Just ("", (), "")) ] where runParserTest nam p input expected = TestInstance@@ -64,7 +73,7 @@ if (output == expected) then return (Finished Pass) else return (Finished $ TestSuite.Fail- $ show output ++ " ≠ " ++ show expected)+ $ "got " <> show output <> " expected " <> show expected) , name = "parseOnly sepCap " <> nam , tags = [] , options = []@@ -84,7 +93,7 @@ if (output == expected) then return (Finished Pass) else return (Finished $ TestSuite.Fail- $ show output ++ " ≠ " ++ show expected)+ $ "got " <> show output <> " expected " <> show expected) A.Done _i _output -> return (Finished $ TestSuite.Fail $ "Should ask for more input") A.Done _i _output -> return (Finished $ TestSuite.Fail $ "Should ask for more input") , name = "parse Partial sepCap " <> nam@@ -99,8 +108,21 @@ if (output == expected) then return (Finished Pass) else return (Finished $ TestSuite.Fail- $ show output ++ " ≠ " ++ show expected)+ $ "got " <> show output <> " expected " <> show expected) , name = "streamEdit " ++ nam+ , tags = []+ , options = []+ , setOption = \_ _ -> Left "no options supported"+ }++ breakCapTest nam sep input expected = TestInstance+ { run = do+ let output = breakCap sep input+ if (output == expected)+ then return (Finished Pass)+ else return (Finished $ TestSuite.Fail+ $ "got " <> show output <> " expected " <> show expected)+ , name = "breakCap " ++ nam , tags = [] , options = [] , setOption = \_ _ -> Left "no options supported"