Ben-Gurion U. developing new computer techniques to analyze historic Hebrew and Arabic documents
Gurion University of the Negev (BGU) will combine the scientific and scholarly expertise of their humanities and computer science experts in a new project to analyze degraded Hebrew documents. The effort to develop new computer algorithms combines BGU's scientific expertise in computer vision, computer graphics, image processing and computational geometry with the scholarly expertise of historians and liturgy scholars to provide valuable answers regarding Jewish liturgical texts and Arabic historical texts that advance scholarship in these fields.
The technical goal of the research is to develop new state of the art algorithms for analyzing text and combine them into an easy to operate, open source system of tools to aid historical document research throughout the world.
Experiments are being conducted on degraded documents from sources such as the Cairo Geniza, copies of which are located at the national liturgy project at BGU, the El-Aqsa manuscript library in Jerusalem and the Al-Azar manuscript library in Cairo. Most fragments that have been discovered at the Geniza are now in libraries at Cambridge and Oxford universities, the Jewish Theological Seminary in New York, The British Library and in Israel and Paris.
Until now the documents have not been researched systematically. Prof. Uri Ehrlich of the Goldstein-Goren Department of Jewish Thought is the head of the Prayer Research Project at BGU. He explains that, "There was one book that was originally used as a Hebrew prayer book from the 12th century, but had been scratched off, and the parchment used to write an Arabic text (called a palimpsest). Our aim was to read the first book and not the second book. So we needed to find out how the Arab book could disappear and would leave only the Hebrew letters of the original book. This is why the computer sciences and humanities departments at BGU decided to collaborate."
"To solve the problem, we created an algorithm to cover the text in a dark grey color, which then highlights lighter colored pixels as background space and identifies the darker pixels as outlining the original Hebrew lettering," said Prof. Klara Kedem of the Department of Computer Sciences and one of the system's creators.
Many of the new methods will apply to other languages as well, including binarization of highly degraded documents (converting up to 256 grey colors to black and white to facilitate digitization), segmentation of skewed and curved lines and word spotting in both curved and highly degraded documents. Other algorithms will be more language specific, such as paleographic analysis of Hebrew and Arabic historical documents that will include automatic indexing of document collections, determining authorship, location and date of the documents.
The research is being funded by the Israel Science Foundation (ISF). Prof. Ehrlich and other BGU scholars in the humanities will be among those to evaluate the system to be built by Prof. Klara Kedem and Dr. Jihad El-Sana of the Department of Computer Sciences and Prof. Emeritus Tsiki Dinstein from Electrical Engineering.
The group is part of the emerging global effort to understand, manipulate and archive historical documents so that they are available to researchers in paleography, archaeology and historical research.
Source: American Associates, Ben-Gurion University of the Negev
Related
- Splash, babble, sploosh: Computer algorithm simulates the sound of waterThu, 4 Jun 2009, 15:49:42 EDT
- Human eye inspires advance in computer vision from Boston College researchersThu, 18 Jun 2009, 0:50:45 EDT
- World's biggest computing grid launchedFri, 3 Oct 2008, 14:49:29 EDT
- Good code, bad computations: A computer security gray areaMon, 27 Oct 2008, 15:56:44 EDT
- NIST defining the expanding world of cloud computingThu, 21 May 2009, 11:29:42 EDT
Other sources
- Project to analyze old Hebrew papersfrom UPIMon, 17 Aug 2009, 17:21:35 EDT
- New Computer Techniques Developed To Analyze Historic Hebrew And Arabic Documentsfrom Science DailySun, 16 Aug 2009, 23:21:14 EDT
- Ben-Gurion U. developing new computer techniques to analyze historic Hebrew and Arabic documentsfrom Science BlogFri, 14 Aug 2009, 20:35:24 EDT
- Ben-Gurion U. developing new computer techniques to analyze historic Hebrew and Arabic documentsfrom Science BlogFri, 14 Aug 2009, 14:56:19 EDT
- New computer techniques to analyze historic Hebrew, Arabic documents under developmentfrom PhysorgFri, 14 Aug 2009, 14:42:06 EDT
Latest Science Newsletter
Get the latest and most popular science news articles of the week in your Inbox!Learn more about
Popular science news articles
- Beyond sunlight: Explorers census 17,650 ocean species between edge of darkness and black abyss
- Generating electricity from air flow
- Therapy 32 times more cost effective at increasing happiness than money
- Beyond genomics, biologists and engineers decode the next frontier
- It's a gas: New discovery may lead to heartier, high-yielding plants
- Therapy 32 times more cost effective at increasing happiness than money
- Full recovery now possible for an 'untreatable' mental illness
- Beyond sunlight: Explorers census 17,650 ocean species between edge of darkness and black abyss
- Is global warming unstoppable?
- Polyphenols and polyunsaturated fatty acids boost the birth of new neurons
- New evidence that dark chocolate helps ease emotional stress
- African desert rift confirmed as new ocean in the making
- Scientists discover influenza's Achilles heel: Antioxidants
- Nanoparticles used in common household items caused genetic damage in mice
- Therapy 32 times more cost effective at increasing happiness than money