I did this once by first cleaning up the HTML-files with batch processing the files on command line through the tidy-parser and then by filtering the content via PHP-script which also put it in the database. But this is custom work and was only worth it because of > 9.000 documents. I think you are faster by doing it your way (perhaps first extract the content via your OpenOffice-solution and then convrt it to utf-8 with another tool.
It finally worked. My "tests" showed that it was necessary to use both a "good" editor and starting a new installation process using utf-8 for all database. Doing only one of both it didn´t work...
Thanks for all hints and answers!!!