Showing posts with label split. Show all posts
Showing posts with label split. Show all posts

Tuesday, December 29, 2015

Python split line by tab or multiple spaces

During the process of trying to find the best way to pull out all the tokens from a tab or space delimited line no matter how they are separated, I found an awesome answer but also learned a few things.

It turns out that python has a "regular expression" methods library re. re has a split function and you don't have to give it both tab and space characters, \s+ is all you need The \s apparently covers all white space characters and + includes any combination of them. To wit:

line = 'one\ttwo\t\tthree four \tfive \t\tsix'
re.split('\s+', line)
['one', 'two', 'three', 'four', 'five', 'six']

(source: http://stackoverflow.com/questions/8113782/split-string-on-whitespace-in-python)

Also learned: This link shows that | is used to delimit multiple split options in the same re:
http://stackoverflow.com/questions/4998629/python-split-string-with-multiple-delimiters

Friday, July 5, 2013

Python read file dates and rename files with datestamps

A piece of code that's so small that does so many things. Searches a directory for all files matching a pattern, gets the modification date of each file, and adds that plus a label to the file name. Uses glob to handle the searching, gets a list of matching files (may be empty if no matches, uses getmtime to get the mod time and formats it with time.strftime, uses split to dice up the original file name, and replace to slap on the label and time stamp and rejoin the new file name to the original path.

def RenameFiles(Directory, Label):

SearchString = Directory + '/pattern*'
tekfiles = glob.glob(SearchString)
for f in tekfiles:
create_date = time.strftime("%Y%m%d_%H%M%S",time.localtime(os.path.getmtime(f)))
pieces = f.split('\\')
new_name = pieces[0] + '/' + pieces[1].replace('tek','%s_pattern' % Label).replace('.','_%s.' % create_date)
os.rename(f,new_name)

Some links:

Shows use of os.rename, glob, and picking out pieces of a filename with array indexes (which I didn't use in the final code):
http://stackoverflow.com/questions/2759067/rename-files-in-python

Shows use of getctime and getmtime. It turns out that since my directory is full of copied files, that getmtime is more useful because the ctime of the copied file reflects its copy date (and can thus be later than the mtime):
http://stackoverflow.com/questions/10149994/with-python-how-to-read-the-date-created-of-a-file
http://stackoverflow.com/questions/237079/how-to-get-file-creation-modification-date-times-in-python

It took a bit of work to figure out how to take the result of getmtime and format it into a string. The result in unix epoch seconds had to be converted to a time tuple using time.localtime so that it could be formatted using time.strftime. Some of that is shown in the previous links but I also read these doc pages:
http://docs.python.org/2/library/time.html
http://epydoc.sourceforge.net/stdlib/time-module.html

Shows use of glob. I was originally thinking of making a full-featured search like my own version of 'find', but it turned out that I only needed to search specific directories, so glob did everything I really needed:
http://code.activestate.com/recipes/499305-locating-files-throughout-a-directory-tree/
http://docs.python.org/2/library/glob.html

The link which gave me the genius suggestion of using nested replace() calls to do the two filename modifications simultaneously:
http://stackoverflow.com/questions/8687018/python-string-replace-two-things-at-once


Monday, September 24, 2012

Splitting binary file in two

All answers to this question involved using unix. No surprise.

Early hunts for an answer to this question involved the "split" command. However, split seems to be geared for cutting files into evenly sized pieces only, and not a single arbitrarily placed cut.

That's okay, because "dd" is my new best buddy. It's loaded with arguments, all of which have a lovely archaic format, and can do any number of cuts anywhere. The version of this command that I will henceforth use until the grave is any variation of the following:

for first half:
dd bs=1st location after split count=1 if=infile of=outfile
for last half
dd bs=1st location after split skip=1 if=infile of=outfile

The main hints were found in this forum thread:

http://unix.stackexchange.com/questions/6852/best-way-to-remove-bytes-from-the-start-of-a-file

This other thread is nearly all above my head; neat shmott shtuff, but this where I first got a clue of the existence of dd. I wonder what xxd is?:

http://stackoverflow.com/questions/9451890/how-to-dump-part-of-binary-file