且构网

分享程序员开发的那些事...
且构网 - 分享程序员编程开发的那些事

如何查找和替换文本文件中的多行?

更新时间:2023-02-23 18:02:02

读取写入 一次一行,所以整个事情不是

 #创建一个查找键的字典并替换值
findlines = open('find ('\\\
')
replacelines = open('replace.txt')。read()。split('\\\
')
find_replace = dict(findlines,replacelines))

with open('data.txt')as data:
with open('new_data.txt','w')as new_data:
为数据行:
为find_replace中的键:
如果键入行:
line = line.replace(key,find_replace [key])
new_data.write (line)

编辑:我将代码更改为 read ('\\\
')
而不是 readliens() so \\\
isn T包括在查找和替换字符串中

I am running Python 2.7.

I have three text files: data.txt, find.txt, and replace.txt. Now, find.txt contains several lines that I want to search for in data.txt and replace that section with the content in replace.txt. Here is a simple example:

data.txt

pumpkin
apple
banana
cherry
himalaya
skeleton
apple
banana
cherry
watermelon
fruit

find.txt

apple
banana
cherry

replace.txt

1
2
3

So, in the above example, I want to search for all occurences of apple, banana, and cherry in the data and replace those lines with 1,2,3.

I am having some trouble with the right approach to this as my data.txt is about 1MB so I want to be as efficient as possible. One dumb way is to concatenate everything into one long string and use replace, and then output to a new text file so all the line breaks will be restored.

import re

data = open("data.txt", 'r')
find = open("find.txt", 'r')
replace = open("replace.txt", 'r')

data_str = ""
find_str = ""
replace_str = "" 

for line in data: # concatenate it into one long string
    data_str += line

for line in find: # concatenate it into one long string
    find_str += line

for line in replace: 
    replace_str += line


new_data = data_str.replace(find, replace)
new_file = open("new_data.txt", "w")
new_file.write(new_data)

But this seems so convoluted and inefficient for a large data file like mine. Also, the replace function appears to be deprecated so that's not good.

Another way is to step through the lines and keep a track of which line you found a match.

Something like this:

location = 0

LOOP1: 
for find_line in find:
    for i, data_line in enumerate(data).startingAtLine(location):
        if find_line == data_line:
            location = i # found possibility

for idx in range(NUMBER_LINES_IN_FIND):
    if find_line[idx] != data_line[idx+location]  # compare line by line
        #if the subsequent lines don't match, then go back and search again
        goto LOOP1

Not fully formed code, I know. I don't even know if it's possible to search through a file from a certain line on or between certain lines but again, I'm just a bit confused in the logic of it all. What is the best way to do this?

Thanks!

If the file is large, you want to read and write one line at a time, so the whole thing isn't loaded into memory at once.

# create a dict of find keys and replace values
findlines = open('find.txt').read().split('\n')
replacelines = open('replace.txt').read().split('\n')
find_replace = dict(zip(findlines, replacelines))

with open('data.txt') as data:
    with open('new_data.txt', 'w') as new_data:
        for line in data:
            for key in find_replace:
                if key in line:
                    line = line.replace(key, find_replace[key])
            new_data.write(line)

Edit: I changed the code to read().split('\n') instead of readliens() so \n isn't included in the find and replace strings