How to input a regex in string.replace in python?

Question

I need some help on declaring some regex. my inputs as as such

this is a paragraph with<[1> in between</[1> and then there are cases ... where the<[99> number ranges from 1-100</[99>. 
and there are many other lines in the txt files
with<[3> such tags </[3>

the required output should be

this is a paragraph with in between and then there are cases ... where the number ranges from 1-100. 
and there are many other lines in the txt files
with such tags

I've tried the code

#!/usr/bin/python
import os, sys, re, glob
for infile in glob.glob(os.path.join(os.getcwd(), '*.txt')):
    for line in reader: 
        line2 = line.replace('<[1> ', '')
        line = line2.replace('</[1> ', '')
        line2 = line.replace('<[1>', '')
        line = line2.replace('</[1>', '')

        print line

i've also tried, but seems like i'm using the wrong regex syntax:

    line2 = line.replace('<[*> ', '')
    line = line2.replace('</[*> ', '')
    line2 = line.replace('<[*>', '')
    line = line2.replace('</[*>', '')

Can someone help me with this? i dont want to hard code the replace from 1 to 99. thanks

ridgerunner · Accepted Answer · 2011-04-14 15:33:07Z

This tested snippet should do it:

import re
line = re.sub(r"</?\[\d+>", "", line)

Edit: Here's a commented version explaining how it works:

line = re.sub(r"""
  (?x) # Use free-spacing mode.
  <    # Match a literal '<'
  /?   # Optionally match a '/'
  \[   # Match a literal '['
  \d+  # Match one or more digits
  >    # Match a literal '>'
  """, "", line)

Regexes are fun! But I would strongly recommend spending an hour or two studying the basics. For starters, you need to learn which characters are special: "metacharacters" which need to be escaped (i.e. with a backslash placed in front - and the rules are different inside and outside character classes.) There is an excellent online tutorial at: www.regular-expressions.info. The time you spend there will pay for itself many times over. Happy regexing!

yep it works!! thanks but can you explain the regex in brief?

Ignacio Vazquez-Abrams · Answer 2 · 2011-04-14 04:00:53Z

up vote 4 down vote

str.replace() does fixed replacements. Use re.sub() instead.

answered Apr 14 '11 at 4:00

Ignacio Vazquez-Abrams
227k18254450

1

Also worth noting that your pattern should look something like "</{0-1}\d{1-2}>" or whatever variant of regexp notation python uses. – bdares Apr 14 '11 at 4:05

kurumi · Answer 3 · 2011-04-14 04:06:17Z

don't have to use regular expression (for your sample string)

>>> s
'this is a paragraph with<[1> in between</[1> and then there are cases ... where the<[99> number ranges from 1-100</[99>. \nand there are many other lines in the txt files\nwith<[3> such tags </[3>\n'

>>> for w in s.split(">"):
...   if "<" in w:
...      print w.split("<")[0]
...
this is a paragraph with
 in between
 and then there are cases ... where the
 number ranges from 1-100
.
and there are many other lines in the txt files
with
 such tags

asked	1 year ago
viewed	4038 times
active	1 year ago

How to input a regex in string.replace in python?

3 Answers

Your Answer

Not the answer you're looking for? Browse other questions tagged python regex string replace or ask your own question.

Hello World!

Community Bulletin

How to input a regex in string.replace in python?

3 Answers

Your Answer

Not the answer you're looking for? Browse other questions tagged python regex string replace or ask your own question.

Hello World!

Community Bulletin

Related