Book / Document Digital Scanner

09 Oct 2009 09:34 #11 by AiXAdmin
Replied by AiXAdmin on topic Book / Document Digital Scanner
I use OCR a lot, but recognise it's not yet a perfect technology where complex formatting is required.

But for text, paragraphing and tables with good preparation in most cases it does everything I ask of it.

In the past I`ve used commercial packages like Abby - but now I simply scan pages to jpg, clean them up a bit if necessary in a graphics application - bundle them into Acrobat and press the OCR button.

Acrobat will process them accurately and rapidly (producing searchable pages as it goes along) - I then either copy paste into a text editor or export to whatever.

There will normally be a little editing to do, but compared to (my rubbish) copy typing there's no contest.

Also the voice recognition bundled with Vista is very capable if you're prepared to spend some time "training" it.

There's many ways of doing this stuff effectively on a shoestring according to requirements - but if time is at an absolute premium and accurate out of the box whole page formatting is crucial be prepared to spend big.

Interesting topic though...

Cheers
Dave

Please Log in or Create an account to join the conversation.

10 Oct 2009 12:05 #12 by OneEighthBit
Replied by OneEighthBit on topic Book / Document Digital Scanner
You should search you tube for "Book Scanner" lots of videos of the type of kit they use for Google Books.

Theres a few pieces of free software for correcting book scans and taking out the black borders, skew and distortion such as unpaper . I've used it myself a few times. It works pretty well with photographed pages too.

Regarding OCR, a lot of it comes down to how "clean" your scan is and the resolution. If your trying to extract the whole text then that is something that needs manual correction but you can always do what Google Books and Flight Globals archive do - keep the scanned document as the original in a PDF/image format and use the OCR text as search data for the document. In that case the odd wrong character doesn't matter so much and in 99% of cases it makes the original scan of the page searchable.

Please Log in or Create an account to join the conversation.

11 Oct 2009 09:27 #13 by mawganmad
Replied by mawganmad on topic Book / Document Digital Scanner
Actually One Eighth, you steer the thread back to the course I was hoping for, the use and access of such a machine by researchers.
Just think if all released TNA, IWM, AHB, and RAFM (and many others)documents were scanned in using this device and it was made available on-line for anyone to download.

I think that the Flight Global on-line resource is amazing, and the copies are just what is required, no need for cleaning up etc.

James Thomas

Please Log in or Create an account to join the conversation.

11 Oct 2009 09:48 #14 by Peter Kirk
Replied by Peter Kirk on topic Book / Document Digital Scanner
If we are talking PDF output as per Flight, then there is an overhead in respect of files size. Good quality scans take up a lot of space whereas with converted to text only picture and drawings would hog space, the text being very small compared to its picture form.

Books that have no illustrations would be very small in file size if OCR'd and can be zoomed in without loss of quality.

A dilemma for archives everywhere.

File size these days seems to be less of a problem but I suspect in decades forward converting zillions of files to the "new" formats might prove to be the end of some archives.

Microfiche (or similar) is about as future proof as you can get. All you need is a light and a lens (providing the actual film lasts).

No Amount Of Evidence Will Ever Persuade An Idiot (probably not Mark Twain)

Please Log in or Create an account to join the conversation.

11 Oct 2009 11:06 #15 by OneEighthBit
Replied by OneEighthBit on topic Book / Document Digital Scanner
PNK,

You touch on some important points.

Yes, using an original scan as a digital document does have a file size overhead but these days that is only an issue depending on your application.

Storage costs peanuts these days with 1.5 terrabyte drives coming in at around £120. To put that into perspective. If you take a worst case scenario where a dirty, picture heavy PDF document with an embedded OCR search index coming in at 512KB you could fit over 16 million PDFs into 1TB of space or about 0.0007p per file.

If the original scans are done well, that is scanned at a high resolution in just duo-tone with a correctly calibrated white point the plain text pages compress extremely well. The definite black/white split and the higher resolution makes life easier for OCR too. The killer with OCR is using too low a resolution where a flaw in the paper can make it look like a rouge comma, apostrophe or change and L to an R - 300dpi+ is the only way to go in this case. Basically its a question of a little effort making a lot of difference and the large commercial book scanners use all of the above tricks internally.

Regarding file formats, this is something I alluded to when I mentioned the BBC Domesday Project in another thread. You have to use common well documented (even ISO spec'd) file formats for any sort of longevity. PDF format is and open specification so is a good candidate as it is a useful document format, can be made searchable and reproduces well. The alternative is to use the TIFF format for page scans as it's been used for that exact purpose for years.

I firmly believe that digitising archives is the way to go because it means that the valuable (and dare say more durable) originals can be locked up safe where they wont be prone to damage from acidic sweat and mishandling. On a recent trip to the Museum of Army Flying archives I was horrified to see that a rare and original diary of very important historical figure had been ripped and torn beyond recovery simply by people carelessly stuffing it back into the archive box.

I think a lot of archives are rather overwhelmed by the technology out there for cataloguing of information and ultimately shy away. Personally I think it's a case of trying to do to much and having too lofty an ideal. The key is getting the data digitised first. The cataloguing and indexing is the second step.

Please Log in or Create an account to join the conversation.

11 Oct 2009 11:52 #16 by contentd
Replied by contentd on topic Book / Document Digital Scanner
PNK - OneEighthBit - Excellent posts.. thanks

The key is getting the data digitised first. The cataloguing and indexing is the second step.


I sort of agree with this, and from a conservation point of view I suppose I wholly agree.

My one caveat is that the knowledge within the document/items (plus added value from individuals) is equally valuable and of much more practical use. - Just digitising without simultaneous teams engaged in "cataloguing/indexing) runs the risk of creating yet another digital dump.

In point of fact given the extent and nature of our brief we can create the core catalogue and index structure (knowledge base) before anything is digitised at all. - Thus allowing "less skilled" labour to easily participate in this heavyweight and immensely time consuming task.

Carts horses, eggs chickens?????

Please Log in or Create an account to join the conversation.

11 Oct 2009 13:55 #17 by OneEighthBit
Replied by OneEighthBit on topic Book / Document Digital Scanner
Contentd,

I whole heartedly agree and that's the next stage obviously.

I think what many archives/museums believe is that they will have to pay or commit staff/time themselves to sit and index/catalogue everything themselves which is actually the least intuitive way of solving the problem.

A good solution is to use user generated content - i.e. have the people accessing the files/archives describe and label what's actually in it. The PRO does this through its Your Archives initiative. It's a simple Wiki set-up and when searching the catalogue clicking the link at the bottom for a given file will take you to it's wiki page. Here researchers can contribute their own description/information as to what is in the files. Considering the often scant clues in the document titles it often provides much more of a clue as to what's inside and if it's going to be relevant.

If you combine a proper indexing system for assets, user/archivist generated metadata to assist in item description/content and *then* store the actual documents as PDFs with OCR generated searchable content you have a pretty fantastic way of finding what you want while retaining the visual appearance of the originals.

One point, while OCR is great for books I think in the case of things like ORBs and reports you sometimes need to see the original because there can also be a lot of, dare I say, forensic information. Often handwritten notes, corrections, and even pen strokes can reveal a lot more than you can think.

Case in point, I had to do some research on the pilots and load of a particular glider for Operation Market. We found the original load-list but it had been corrected and amended so many times for the planned Linnet/Comet/Market operations it was hard to tell what corrections belonged to what. In the end we figured it out by matching the various pen stroke widths and writing styles against the operation names to figure out which belonged to that. If we'd just had the raw text to go on we would of never of figured it out. :)

Please Log in or Create an account to join the conversation.

11 Oct 2009 16:44 #18 by AiXAdmin
Replied by AiXAdmin on topic Book / Document Digital Scanner
OneEighthBit - Thanks - another interesting post.

Hopefully this topic will keep running and gain momentum - as the views of each end every member of this board are of real interest.

Cheers
Dave

Please Log in or Create an account to join the conversation.

11 Oct 2009 18:46 #19 by Peter Kirk
Replied by Peter Kirk on topic Book / Document Digital Scanner
Yes, OCR can only be applied to printed documents and possibly type written (with all theproblems that entails).
Archives like Flight create a bit of a dilemma for me. On the one hand it is good to see the original layout and print but on the other it would be good if it was clearer and zoomable with clarity (pictures and drawings excepted). On the other hand maybe I should go to Specsavers.?

No Amount Of Evidence Will Ever Persuade An Idiot (probably not Mark Twain)

Please Log in or Create an account to join the conversation.

11 Oct 2009 22:47 #20 by OneEighthBit
Replied by OneEighthBit on topic Book / Document Digital Scanner
PNK,

I believe the problem with the Flight Global archive because of the way they scanned them and the particular software they used.

First of all I have to applaud them for what they have and for putting it on-line for free and I think it shows great initiative. From the quality of the scans and the errors that creep in (remember the woman's hand including wedding ring that covered one page?) it's obvious they undertook the whole thing themselves using a simple desktop scanner and Abbyy FineReader to OCR read the pages and produce the searchable text.

The main issue is that most desktop scanners aren't big enough to happy get anything much larger than A4 comfortably only the scanning bed and hold it flat which is why you tend to get the shadows/distortion towards the gutter. Also, with older magazines you get yellowing of the pages which comes out as grey and makes life harder to the OCR software and also adds to the overall file size due to the way the compression works. For pages with just text on them scanning them as higher dpi duo-tone is a solution but I suspect the scanning software noticed the artwork/photographs and kicked back to greyscale scanning mode. I also suspect it used a default low resolution of about 72 to 100 dpi for the scans which is why you can't enlarge them.

It's possible that if they'd changed a few settings in the software they might of got better results but that we'll never know. Either way, they probably got the results they want which was an online low-res but readable archive of their past magazines.

It is possible to create print quality PDFs and in fact the print industry now does actually use such things when sending jobs from agencies to the print shops. Obviously to achieve that you have to put the effort in. Once I spent about 16 hours in total scanning and re-touching a 100 page Air Publication to be made into a print quality digital PDF. I got the results I wanted but I had to put a LOT of effort into it :D

Please Log in or Create an account to join the conversation.

Time to create page: 0.045 seconds
Powered by Kunena Forum

We use cookies to improve our website and your experience when using it. Cookies used for the essential operation of this site have already been set. By continuing to use this site you are agreeing to this. To find out more about the cookies we use and how to delete them, see our privacy policy.

  
EU Cookie Directive Module Information