Batteries
I forgot to mention in my last post that I went and bought a couple of batteries for the PowerBook. Three batteries should give me a good chunk of my upcoming trip with access to the notebook. I'm hoping to get some work done (on about three separate projects) as well as get in some movie time. :-) I have a couple of new DVDs that I've wanted to put some time aside for. I'm guessing that watching a DivX is a lot less power-hungry than spinning the drive of a DVD, so I'll convert them to DivX first. Kind of like ripping a CD for an iPod: morally it's just fine, as all I'm doing it transferring from one media to another, but legally you can't do it (at least, not in Australia. No "fair use" here). Then again, legally I can't listen to my CD's on an iPod either. :-)
Other than movies, I'm hoping to write a few more notes on the structure of the data files in Kowari. Some of this is obscure, and is really only documented in the class implementations (where "documented" means that the code tells you how it's done). I understand it in a general sense, and the code helps clear up anything I forget, but it will be very useful to have real documents that explain it. I've already started on this with the Free Lists (probably the hardest format to follow). This slowed down the Javadoc work, so I'll probably put aside the file format until the flight so I can get through Javadoc before I leave.
Other than that, I still have OWL code to write (I want to put together some subsumption code soon), and I'd love to do a bit more on jCarbonMetadata, which I haven't touched for weeks. Speaking of which, I started on an objective-C tutorial and meant to get back to it. After all, instead of porting the missing Carbon functionality to Java, why not learn objective-C and write a bridge from Cocoa to Carbon? :-)
Yes, this is a lot to do on one flight... but I'll be traveling from 11am in Brisbane until 5pm in D.C., all in one day. That's a total travel time of 20 hours. And no, I can't sleep effectively on flights. My plan is to keep myself busy, which is why I need several things to keep my attention (so I can swap when I lose concentration).
Huxley
The other thing I'm planning on doing on the flight is continuing the novel I've just picked up. I'm finally reading Aldous Huxley's A Brave New World. It's always strange to read old science fiction, as the authors had no idea how far we would come in the twentieth century (computers weren't even invented for decades after Huxley wrote this novel). It's easy to get distracted by thinking different the real world is to the story, so it can take some real concentration to see the point the author was trying to make. However, I'm only into the second chapter and already I love it.
I particularly love the noise-aversion and shock therapy given to Delta babies to make them hate nature and books. It turned my stomach, which was obviously the intent. But the reasoning was masterful. Discouraging Deltas from reading prevent them from wasting time on something they didn't need (something logical for this future world). But it was the reason for hating nature that really struck a chord for me.
Deltas were once allowed to enjoy nature, as they would then consume transport, but that's all they consumed. Nature was gratuitous, and kept no factories busy. So it was decided to abolish a love of nature. However, the desire remained to have Deltas go out to the country to consume transport, even though they now hated the country. So Deltas were made to love country sports, which required the use of expensive and elaborate equipment (thus creating more consumerism). Beautiful.
The scary part about this concept is that modern society mirrors it so much. Sure, we don't condition children with shock therapy... we use advertising instead. However, the end goals are the same, and in my opinion, the fact that we do it to children is just as deplorable.
Huxley wrote this book as a warning to the world about the direction we appeared to be going in. It has had enough influence to slow down some development over the years, but it never really stopped it. It's obviously an insightful story, and I'm looking forward to reading more of it.
Chicago
Advertising at children brings me to a new point. When Luc wants to watch TV we let him see ABC-Kids. This is a brilliant service, as it is government funded and has no commercials. I don't want him to watch too much TV, but at least there is something he can watch that I approve of (the kid is a "Wiggles" addict). I'm wondering if there is a similar TV station in the US? From what I've seen, I really doubt it.
This is just one of many questions I've been pondering lately. I'm considering a move to Chicago to take on a new job. Three years ago I'd have agreed at the drop of a hat, but now that Anne and I have a family it changes the rules. We have a great lifestyle here, and there seem to be some good opportunities coming up if we stay here. It also looks like it would take a bit of money to move to the States (on top of what we've been offered to help us move), and we don't really have any at the moment.
From what I've learned about Chicago, we'd be taking a backward step in our lifestyle (based on what I'd be earning). On the other hand, I've been wanting the experience of working overseas, and this could be the opportunity to do it. Besides, even if we take a backward step, then we can still take forward steps in the future.
The other thing is that (almost) everyone I've spoken to about Chicago has recommended it as a great city to be in (though I haven't heard any specific reasons why). Alternatively, Anne is scared about managing two kids (including a newborn) in a city where she knows no-one. So it's a bit of a toss-up for us. I'll be going to Chicago at the end of my trip to have a look around. Hopefully that will help us to make the decision (though I'll need Anne's support).
I'm mentioning it here in case anyone knows Chicago (and the US in general) and has an opinion to share. :-)
Friday, July 01, 2005
Thursday, June 30, 2005
Axioms
Well the test didn't work, but I had to go to bed. When I got up in the morning yesterday, I spent 15 minutes looking at the problem, and discovered it wasn't really a problem. The code was working just fine, but I had a typo in the RDF for the axioms. Doh! :-)
It's really impressive how many entailments come out of the axioms alone. There are 33 axioms and this was increases to 120 statements after the entailments. When I ran the RDFS rules over the RDFS model (and it's OWL description) then it turned 643 RDF statements into over 1259 statements.
It was while I was looking at this that I discovered how useful entailments are, even simple ones like this. For a start, all of the types suddenly become available, even those that are not immediately evident. This in turn showed up a couple of simple errors I had in the RDF which had slipped past me before. It turns out that entailments from bad data can lead to a lot of statements that are very obviously wrong. Quite handy.
I didn't get to work much more on this, as I needed to get back to the NGC preparation.
XA Store
Yesterday I was lucky to have DavidM online while writing Javadoc for the transaction layer. This helped me to learn a lot about the newer parts of the code much faster than if I'd had to do it on my own. This will definitely help with the handover.
Today I spent several hours working out a timetable for the handover. I expected to do this in a couple of days, but DavidW needed something from me quickly. So I tried to write an email with a broad outline. However, I ended up writing a lot more than I thought I would, and I think I'm halfway to the full preparation. :-)
One more thing I'd like to do is to write up some diagrams of the file formats. This is all available by reading the source files, but they have to be carefully interpreted to work it out. Having diagrams on hand should be very useful.
Well I'm exhausted tonight, so I'll leave it here and get to more Javadoc in the morning.
Posted by
Paula
at
Thursday, June 30, 2005
0
comments
Tuesday, June 28, 2005
TKS
For anyone who missed it, the rights to TKS have been purchased by Northrop Grumman. Importantly, they've agreed to support the Kowari project. Wow.
I've agreed to do a handover of the code, and as a part of this I'm writing Javadoc for the XA store. I helped write this over 4 years ago, so I sort of know it, and I sort of don't. DavidM has modified a few things since I was last in that code, so it's been a learning experience for me. Fortunately the process of Javadoc is helping me learn it all.
I'll be flying to the US next week on the 6th of July, and will be returning on the 27th of July. I'm not looking forward to missing Luc (and Anne). On the other hand, I love traveling for work, particularly to the States, so it won't be all bad. It will definitely slow my blogging down though.
Rules
In the meantime I'm still working on the rules engine and the tests. It appears to work remarkably well. Must be all that effort I put into it. :-)
Since I have to work all day for NGC at the moment, I'm spending my evenings (and the upcoming weekend) working on the rules. All that is left of my paid work is the proper set of tests, and perhaps a short document describing how to write a rule configuration (in case something more than RDFS is desired). I'm pretty sure I can get all that done before I fly out.
Many of the tests specified by the W3C are quite trivial, but they reminded me that RDFS requires axiomatic statements. That is fine, as I could just load them into the destination as a separate RDF file. This bothered me though, as it's a very manual process, and a necessary one for many rule systems. So rather than expanding on the tests tonight I've been encoding the axioms into the rules engine. Now the RDFS rules file includes all the axioms, and the rules engine knows how to read them in and insert them at the beginning of execution. I still need to test that this is working correctly, and it is compiling as I type.
It's now 1am, and I have to give Luc his bottle when he wakes up in the morning, so I'd better run this thing and get to bed (please work first time!)
Posted by
Paula
at
Tuesday, June 28, 2005
1 comments
Monday, June 27, 2005
Tests and Improvements
The tests for the rules seem to be going fine. I could continue to write tests for it forever, with each one getting more complex, so it's hard to know how far to go. After all, they can take a long time to write. At the moment, I'm looking to test each RDFS entailment one at a time. Eventually I'll be moving on to how entailments interact, but I'll only go so far with that, as I really want to move on to the next part, which is OWL entailment.
Still, the code (without tests) is checked in. Have a look in the "rules" directory for an example of running the RDFS rules against the RDF for the RDFS rules. ;-)
I had a talk with Andrae today about how RDFS/OWL and how the Rules engine in Kowari works with it. I also showed him some of the internal structure, and he had a couple of ideas that will help with efficiency. I'm starting to accumulate a list. It's frustrating, as I want to make it all as fast as I can, but I also want to make it as functional as I can, and I can only work on one objective at a time. :-(
The efficiency things I want to do are:
- Make the various implementations of
Tuples.getRowCount()run in log time, rather than linear time (this needs a count in the nodes of the AVL tree). - Add a new version of
Session.insert()which accepts anAnswerrather than aquery - Properly merge the source and destination models for rules if they are the same model
- Keep a map of
Constraints to their counts for each rule, and ensure that answers are using theseConstraints.
Javadoc
I've also been adding some long-needed documentation to the storage layer.
I pair programmed with with DavidM starting 4 years ago. I worked with him on it for over a year before moving onto other things. So I mostly know how it works, but there are subtle changes since I saw it last. Plus it's been nearly 3 years since I last worked in that area. Back when this code was written, David and I were really pushed for time. To help us get it done in time we were told that we could skip the documentation. It allowed us to meet the deadline that we had, but of course we were never given the time to go back and document it like we were supposed to.
It's been an interesting experience coming back to something so complex after so long. I don't remember all of it clearly, but this process of documenting the methods is really helping me work it out. It's also interesting discovering some of the changes which have come about since I last worked on it.
Posted by
Paula
at
Monday, June 27, 2005
0
comments
Saturday, June 25, 2005
Whew
The rules engine works. Yay. :-)
While thinking about the RMI problem on Thursday night, I realised that I do want the rule configuration to serialize (I use the American spelling, with the "z" here, since all the interface and method names are spelt that way). The reason for this is because the server with the rules might be different to the server with the data.
My concern with serializing like this is unnecessarily serializing from the server to the client and back again, particularly when the client is on a separate machine to the server. Serendipitously, I'd already wrapped the rule structure in a remoteable class. This works well, because nothing gets moved to the client at all. When the server gets this reference and asks for the rules then this is the only time that the rules get serialized. If the two servers are separate then the RMI serialization is necessary, and this approach works just fine. If the server for the rules is the same as the server holding the data, then the rules will still be serialized, but it only happens once, and the RMI transfer is within the one JVM, so it should be very fast. Either way, there is no unnecessary transfer to the client.
It's amazingly satisfying to see the entailments being generated. :-)
Merging
For stability purposes, I've been coding against a version of Kowari that I picked up a couple of months ago. My code has been largely independent of other modification in the system, so this should be safe. Nevertheless, I'm about to update everything and confirm that it all still works.
Once that's done, I'll need to tidy up my tests in order to check everything in. Since it's all working (and can't break any other components), I might check it in without all of the tests. That's because a few people have been asking after these rules. Hopefully they'll like them. :-)
The algorithm for running the rules is not quite complete. At the moment each rule is dependent on finding a change in the total result of a query in order to determine if it needs to run. This requires processing, even when the rule does not need to run. The completed algorithm tests each individual constraint for changes, before testing the joined result. This means that there will be no join processing for (almost all) rules that don't need to be run, and no additional expense for those rules which do need to be run. I expect these changes to take just a few days, but they are not needed straight away.
One feature that might be nice to add would be to keep a total of all entailed statements and return this to the user when done. I'll look at doing this soon as well.
Eclipse
I've gained quite a bit out of using Eclipse, but it's been frustrating me lately. I often find myself typing up to a line ahead of the rendering, and often have to wait for the IDE to finish processing before I can do anything. As a result, I've started using VIM again whenever I want to do fast changes. I'll have to work out some way to use Eclipse again without having to spend half my day waiting for it.
Maybe it's just that my notebook is too slow. It's only a 1.33 GHz G4 (with a GB of RAM). This isn't as fast as many desktop systems, but it's quite responsive on every other application I use, so I expected that it should be fast enough.
Last night I had a go at running Eclipse on my Linux desktop, and piping the display to the X server on this notebook, but that kept crashing. On the Linux-GTK setup it crashed with a GTK Window error whenever I selected the "File" menu. So I tried the Linux-Motif setup, but that caused a Hotspot error. Yuck. I haven't tried it on the local display, plus I was using Java 1.5, so there are a few things that could be causing these errors. I guess that will keep me occupied for a day or so, when I get the time to look at it.
Posted by
Paula
at
Saturday, June 25, 2005
2
comments
Thursday, June 23, 2005
Slow and Steady
I've been plodding along with the rules engine over the last week. Unfortunately, I'd initially planned to be done by now (not planning on getting sick, etc), and so I'd already committed myself for some other modeling work. That meant that I had to juggle the two together. However, I really need to finish the rules work quickly. I want to get it working, so I can move on to the next part of OWL support (needed for my thesis) and also because I won't get any money for my last couple of months until it's done. I think I'm actually supposed to invoice for time rather than the product, but I made a commitment to complete this phase, so I'll have it working before sending in the invoice. I just hope that the payment won't be too long after! :-)
After this last week, the RDF seems to be configured correctly, the code reads it all as required, and the elements all work properly. But it's tough getting them to work together in the way that I want. I'm half tempted to glue it all together in a way that I've already seen work, but I need the more general solution. At the moment, the problem is RMI. That figures.
When I tell ItqlInterpreter to apply a set of rules, I want to read the rules structure and return a result to the ItqlInterpreter. Then ItqlInterpreter can change it's current session to point to the data to be worked on, and run the rules. This all seems to work (it's not fully tested, but what I've done so far works) if I have it all occur in one step on the serve, but then the rules can't be separated from the data to be worked on (though I am presuming it is all found on the same server). I need to make it happen in the two steps I've outlined.
Unfortunately, I'm having real trouble getting the rules structure to pass over RMI correctly. By default, getting the rule structure from ItqlInterpreter would serialize the structure and create it in the client space. This has two problems. The first is that it's inefficient. There's no reason to move all that data over to the client, since it will only ever get used at the server. The more important problem is that the client would need to have access to the rule structure classes for de-serialization, and I'm trying to keep the client completely oblivious to all but the interfaces.
The better way to manage this situation is to pass back a remote reference to the class. I struggled for some time getting this right (I forgot how annoying RMI can be in this regard). To simplify things I decided not to ship a reference to the entire object (since the methods on that object should never be called remotely), and instead created a remotable wrapper to hold a local reference. The remote reference to the wrapper can be shipped over RMI, and when it comes back it can be queried for the local object that I want. Only that's not what I'm getting.
It took a LONG time to get a stack trace that was useful (it's being completely hidden in RMI, and was never being thrown from where I thought it was being thrown), but I finally worked out that the problem is coming from the server when it tries to extract the local reference to the rules structure. At this point it tries to serialize the rules structure, which is not legal (intentionally so). I believe that this is because the object doesn't know that it is now back on its machine of origin, and does not know that it can just pass back a local reference.
Perhaps I'm approaching it all wrong. Now that I have the RMI compiling and running correctly (this took me some time) I could possibly drop the indirection of the wrapping object, and pass back the rule structure as a remote reference (like I originally tried). My only concern is that calling run() needs to pass in a DatabaseSession, and I want it to take the local session without trying to serialize it. I'll have a go in the morning anyway, and see what happens. In the worst case I can always keep a local map of the remotable objects to the local rules structures they represent. Then when an object comes in for a "run request" I can get the local rules out of the map.
It's late. I'll think about this overnight and have a go in the morning.
OWL
Bob asked that I work with him on looking at extensions to OWL. He knows I've been thinking about this lately, and it turns out that he has some ideas as well.
Yesterday was the first opportunity I've had to really see how he works on these problems, and I have to say that I'm impressed. While he doesn't have a strong background in OWL, he is able to draw on a good formal understanding of category theory, E-R diagrams, and other areas. It allows him to see where OWL is missing functionality, or how certain functionality might be achieved using OWL constructs. There were a few occasions when I could suggest using an OWL or RDF construct to achieve something, and he could quickly show why this would or wouldn't work, and if it wouldn't work then why not.
The most impressive thing is that he really knows the boundaries of his knowledge, and he knows where to go looking when he doesn't know something. Conversely, I don't know how much I don't know,. Even when I recognise that I need to learn more about something, I don't even know the name of the field that I need to learn more about. I guess that just comes with experience.
The two big things to come out of our conversation was Cartesian product classes, and predicate composition. Cross products allow for several important relationships, but most importantly they would permit the description of a pair of relationships which may be individually repeated, but together must be unique (like composite keys in a database table). Predicate composition ends up covering a couple of the ideas I've already described, such as Euclidean predicate relationships, only it does it in a single construct.
I know that Ian Horrocks has explained that things like predicate composition can lead to undecidability, which needs to be considered for OWL DL. While that is no problem for OWL Full, a lot of people are more interested in OWL DL for the moment, so I think that it is important that any new constructs have some use in OWL DL. However, decidability is not something that either of us know a lot about. Fortunately, I've been given a reading list by Guido, plus I have those pointers from Ian. I just need to find some time to sit down and read them! :-)
Posted by
Paula
at
Thursday, June 23, 2005
0
comments
Tuesday, June 14, 2005
Quiet
Just like every night lately, I'm tired... so this will be short.
Testing and debugging went well over the last week. Since I found the dynamic compiling and error reporting of Eclipse to be so useful I decided to sink a bit of time into making the project integrate properly. I made some real progress, but it's still not done completely. Fortunately Eclipse still lets me run programs when errors are present in the build, otherwise the work I'd done would be a waste of time. It would still be nice to get everything working though. In the meantime, I decided I've spent enough time on it for now, else I won't be saving any more time than I've spent.
There are so many details that I learnt in setting this up, and it's annoying that I didn't blog them all as I went... particularly the problems. For instance, CVS didn't pick up one of the directories on SourceForge. Why not? No idea. But I could use the CVS perspective to "Check out" a subdirectory to an already existing project. So I just had to find the missing directories, and give them the correct destination.
The debug environment that I have running now lets me use breakpoints, and lets me properly analyse Kowari while it's running. The last class loader (with embedded jar loading) prevented this from working, so it's a significant step over what I used to have. However, it's still annoying that I can't get Log4J working properly. I've discovered that the correct config file is loading (I put an error into the XML, and looked for an error message), but it ignores the "Category" statements which stipulate that Debug messages should go to the appenders. The debugging environment helps here, but logging is still important (especially when you consider how slow Eclipse is). For the moment I use warning messages instead of debug messages, and I'll have to remember to change them all back before checking in.
Anyway, the configuration now works (of course, the problems were all in those areas I tacked on to the RDF at the last minute), and I've traced all of the problems, which I'm currently working on. It should be finished by the end of the week, but I still need time to write all the tests.
Public Holiday
Yesterday was a public holiday here (the Queen's birthday - not that anyone seemed to care about that). I worked for some of the weekend, and I planned on working for Monday as well, but Anne also needed to work. She's supported me a lot with work lately, so I agreed to look after Luc on my own. I spend a lot of time with him every day, but yesterday was the first time in a long while that I spent an entire day with him, where he was the centre of attention. No internet access, just taking him places and doing things with him. He's doing very well at the moment, trying out new words and solving lots of everyday problems.
We had a great day, and I'm pleased that I was "forced" into it. I'd better concentrate on work for another few weeks, but when it's done I should take another full day like that. It's good for the soul. :-)
Posted by
Paula
at
Tuesday, June 14, 2005
0
comments
Monday, June 06, 2005
Substantive Coding Done
I spent the last few days concentrating on coding rather than blogging, as I have been getting close to the end. I've now finished the substantive coding for the first iteration of the rules engine. That doesn't mean it's done, but I feel good about it anyway. :-)
What I mean by substantive effort is that the code is all written, and it compiles... but that's it. Except for iterative testing during the process of coding, I haven't even tried to run it yet. I plan on starting the debugging tomorrow. This always takes time, but just starting on the debugging stage makes me feel like I'm in the home stretch.
The current implementation is based on performing the equivalent of a "select", followed by an "insert/select" if needed. This is a lot like the original proof-of-concept code, only now it is fully integrated and is being done on the server side. This is not the ideal implementation, but it will work, and should demonstrate that everything is implemented correctly. It should work pretty quickly too.
I think I'm on schedule for completion of this round of the paid work, but a few sick days recently mean that I'll need to go beyond my scheduled finishing date. I believe I'll have everything written that I'm supposed to have written, but in reality I'm being paid for time, not for completed code. The other reason to keep working on this full time is that I'm enjoying it! However, I can only afford to go for the extra time needed to make up those days I was sick. Hopefully the system will be so useful and show so much promise for further development that someone will pay me to do more of it! (Well... I can dream, can't I?) :-)
Once I have it working, and RDFS is executing correctly, I can move on to the next stage. This version will count the size of each individual constraint, rather than the size of a completed Answer, making it much more efficient. More importantly, it will match the design I'm writing my thesis about! The initial code to do this should take less than a day, but I'll need to spend some time in the query engine to make sure that constraints are being cached and re-used correctly. I'm not sure if that time should be considered "coding" or "debugging", as I'll be using constraints in a manner slightly outside of their original design (which feels like re-designing and coding), but I'll be approaching any problems like I would any unexpected error (which is debugging).
Rules vs. Ontologies
A more pressing concern is the need to make transitive constraints accept a variable predicate. This is not needed for RDFS (since the only transitive predicate is rdfs:subClassOf) but it will be needed for OWL. Once OWL is introduced, the specific rule for transitivity of sub-classes can be dropped in lieu of a declaration in OWL that rdfs:subClassOf (and owl:subClassOf) is a transitive property.
This brings me to a point that I've been thinking about for a while. I sort of understood it before, but I think I've only just started to really get it. What does an ontology language give us that we don't get from rules? After all, ontology inferencing (and consistency checking, but I won't go there right now) is performed by rules. Rules also allow much greater flexibility than we can achieve with a ontology languages like OWL. The commonly cited example of OWL's limitations is that it can't express the "uncle" relationship. An "uncle" relationship is relatively straightforward. If person A has parent B, and person B has brother C, then A has an uncle C. This is easy to describe in rules, and impossible in OWL.
If we can do everything in rules, and OWL is limited, then why use OWL?
The answer (for me) is demonstrated with the transitivity of subClassOf. If we just had a rule system, then we would need to have a rule for inferences on this predicate. ie:
if AThat's fine, but what about "less than"? We need a new rule:rdfs:subClassOfB
and Brdfs:subClassOfC
then Ardfs:subClassOfC
if A < BHow about "greater than"? New rule. "Equal to"? New rule. Every time a new transitive predicate appears we need a new rule to handle it. This means that rules have to be very domain specific. They can't handle anything that wasn't known about at the time they were created.
and B < C
then A < C
However, using OWL a predicate can be declared to be an
owl:TransitiveProperty. Suddenly we have just one rule:if property is transitiveWhenever a new property is introduced which is transitive, then we can just declare it in OWL. Of course, this goes for all of the properties of properties that are definable in OWL. So the ability to describe the properties of a property means that we can write generalised rules to make deductions on them.
and A property B
and B property C
then A property C
I've sort of understood this for a while, but it was only while thinking about Euclidean properties that it finally crystallised for me. Ideally, it would be possible to assert something like:
parentOf isEuclideanTo siblingToSo we could end up with a rule like:
if property1 isEuclideanTo property2Of course, isEuclideanTo would not be symmetric, though it would be possible to infer backwards on it (a person's parent must be the same as their sibling's parent).
and A property1 B
and A property1 C
then B property2 C
There are more complex types of relationships between relationships. Uncle is a good one to demonstrate this, as it requires 3 different types of relationship which are all related to each other. While possible, the RDF required to describe something like this is starting to get messy. Complexity is introduced when you realise that one of the relationships can actually be either "brother" or "brother-in-law". Also, an uncle relationship can be deduced, as can a nephew/niece, but the parent in the middle of the relationship cannot be.
All the same this kind of knowledge about properties is something we use every day. If an ontology is to describe real world objects and relationships then it will need to be capable of describing relationships between relationships.
With this in mind, I ask the OWL mailing list what people thought of such a construct. I half expected to be shouted down, but at least I'd get to find out why. Instead the response was encouraging. Ian Horrocks (who wrote half of the papers I cite) explained that property relationships are indeed useful, but that care must be taken to ensure they are decidable. He's suggested I read one of his papers on the topic.
Who knows? Maybe one day OWL can include something like this.
Posted by
Paula
at
Monday, June 06, 2005
0
comments
Friday, June 03, 2005
Modeling
Sure enough, the sore throat developed. I'll sure be glad when my immune system is used to seeing all these bugs that Luc brings home. It slowed me down a little in the last couple of days, but I still wrote a lot of code. Well, I think it was a lot. :-)
The main impact was that I didn't go out to exercise (which I really need to do in order to work efficiently) and I was also too tired to blog. I'm too tired again tonight, but sometimes you have to push the envelope.
The code over the last few days has been building data structures based on RDF. I had already done some of this work before, but it turns out that integration has needed a more thorough exploration of the data store.
I've realised that there are two ways to write this code. I can write it with a full knowledge of the data structures I'm reading, or I can write it with very little knowledge, and use the ontology to tell me how to build the data structure.
I started out by using the ontology of the rules just a little, while mostly relying on my own knowledge of the data structure and putting that explicitly into the code. This is fine, but it's not very extensible. It was while building the constraint tree that I started to see some other potential problems with the explicit approach.
Each node in the constraint tree is either a leaf, or refers to two or more child nodes. To query an RDF structure about a tree of arbitrary depth it is necessary to get all the links from parents to children as a set, and to connect them together. As each node is found for the first time, the associated Java object must be created to go with it. A problem can arise here if the node is of a type that extends another concrete type. In that case it is necessary to put off creating a node until all of its types have been found, and to then build the most specific type. This is where the ontology starts to play a part.
For a start, simple RDF will just give me the concrete types of the nodes, with no information about the superclasses. It is only with an ontology that the other types would become available (I don't know about anyone else, but I'll probably end up running the inferences against the rules themselves, just to see if I can). At that point I'll need to structure the types together (using transitive owl:subClassOf) and check that there are no loops (except the obligatory subClassOf(A,A)), before finding the most specific type to instantiate. Instantiation will be of a class that would be associated with the node (I don't have that yet, but it would be trivial to add). Of course, multiple inheritance makes it that little bit harder.
The advantage of such an approach would be that extra classes in the structure would just require an OWL definition. Then they could be used in the RDF with no changes to the Java code.
This would apply to any kind of data structure that is expressed in RDF, not just my queries. Building a system like this would be similar to a UML modeling tool with runnable objects (something that I know some tools do). It's cheating a little to link it to a class name, but that is where the idea of storing a Java AST (Abstract Syntax Tree) in RDF would come in. In that way, the basic ontology of a set of classes could be written in OWL, with RDF annotations to describe the complete implementation of the classes. An RDF structure with rdf:types referring to these classes would describe an instance graph.
I really like this idea, and it shows some of the modeling power that comes with OWL and RDF. Unfortunately this would take a few weeks, and I don't have the time for it right now. Doing a Java AST representation in RDF would take a lot longer, though the compiler would be fun.
For the moment I'm sticking to what I know of the data structure, and almost ignoring the ontology. That's a shame, but it's letting me finish it in just a few days, instead of weeks. Fortunately the none of the instantiable classes in this system have descendents, so there won't be any potential confusion about the class to instantiate for any nodes in the constraint tree.
Eclipse
I spent Wednesday morning with Brad. He offered to show me how to get Kowari working with Eclipse. I had a go at this last year when we had the jars-inside-jars class loader in Kowari, and configuring this was very difficult, highly manual, and didn't really work well. It seems that the new flattened class arrangement works much better.
Brad was also able to show some of the useful tools included with Eclipse, such as refactoring, which certainly makes it seem quite compelling. I haven't made the transition just yet, as I haven't got CVS going yet. Eclipse does not seem to understand CVS directories that were created outside of it, so I'll need to get a fresh checkout, and port over my modified files. That might be a job for the weekend.
More importantly, it was very good to have a chat with someone who is using Kowari. It helps to keep focus on what I'm doing it for when I speak to people who actually interact with this stuff. In a similar way, I worked with Andrae this morning, and I appreciated hearing what another developer is doing.
Final Word
Oh, and I promised to make a comment about a "strange Luigi guy". If he's reading, then "Ciao".
Posted by
Paula
at
Friday, June 03, 2005
0
comments
Wednesday, June 01, 2005
Modal Logic
Luc got himself kicked out of day care yesterday. He was running a temperature, coughing etc. So I had to babysit while working today... slowing things down quite a bit. That means more work on the weekend, unfortunately. I hate to think what my sore throat tonight means!
Fortunately, my sister was able to watch Luc while I went in for the Logic group that's held on Wednesdays. Today we covered Modal Logic. I found it all interesting, right up to the point where we looked at the group of Kripke logics.
These logics are based on 5 axioms, labeled T, D, 4, B and 5. Each of these axioms describe specific types of relationships: Reflexive, serial, transitive, symmetric, and Euclidean. It was this last one that got my attention.
OWL has no mechanism for describing how relationships relate to one another (other than with inheritance). Consequently, it is impossible to describe an "uncle" relationship as being built from a "parent" and a "brother" relationship. This is often devolved to a rule language to perform. However, it is a valid and useful thing to describe in an ontology.
Euclidean relationships go some way to addressing this. They allow a description between entities A, B and C. If A relates to B and C in the same way, then B relates to C in some other way. That doesn't work for the "uncle" relationship, but it would work for describing siblings if they share a common parent.
I wonder if there is some way to incorporate this effectively into OWL?
Linking Blank Nodes
I finally worked out the best way to traverse these nodes.
The rules all need to sit in memory (even though the corpus of RDF data does not), so it is appropriate to build up the data in memory. I was concerned about the difficulty of this, particularly when it involved new queries for every branch on the tree, but I now realise that this is not needed. Instead, I have been able to get most of the information in a raw form with a simple query, and put it into HashMaps. Then I can walk my way through the data quite easily.
More importantly, this has the added advantage that almost all of the data comes from a single query. This means that there are no concerns about blank nodes comparing incorrectly, regardless of transactions. However, I've tried them from one query to the next, and all seems well. This was important, as I really needed to split the queries up into at least two anyway. My only other option was to union the results of two vastly different queries together, using the same variable names, along with predicates which are given variable names but set using <tucana:is>. Yuck.
That reminds me. No one commented on me wanting to change all the tucana references to kowari. That must mean that no one minds, huh? :-) Perhaps I should change the code to accept either for the time being, before finally dropping the tucana support.
Posted by
Paula
at
Wednesday, June 01, 2005
2
comments
Monday, May 30, 2005
Traversal
It's late, so I'll be brief.
I need to traverse my way down the constraint tree of a where clause. This is an issue because none of the nodes are named. So the only way to travel down is with a set of conjunctions in the iTQL query. Unfortunately, every level down in the tree means a new conjunction, and the tree is arbitrarily deep. So how do I go down?
There are a few solutions.
The first is to name everything. That would work, but should not be necessary, and would be a pain to use. Besides, an unnamed node should not cause it to all stop working.
Another solution is to use JRDF and traverse my way down manually. This is undesirable as I'm already using iTQL heavily. It also runs into a problem of re-using blank nodes from one query to the next. That should be OK, but strictly speaking is not allowed.
An alternative is to automatically generate iTQL to traverse its way down the tree. This would work, but would also lead to messy results with a lot of work to interpret. It would get even harder if different branches on the tree were different depths.
Thinking about the shape of the tree made me realise that the form of the queries is always in conjunctive normal form. I could take advantage of this, but it will lead to a very inflexible system in future... and I may need some flexibility when OWL gets fully implemented. However, it's always a fallback position.
My final option is to use blank nodes from a previous query as constraint elements for a new query. This should work, particularly as I'm on the server, and if I stick to a single transaction. The problem here is that I need to abandon iTQL and build the queries by hand. Fortunately, this isn't too hard, and need only be done for this part of the rule parser. I also intend to use iTQL to pre-query the simple constraints, and hold references to them for use as I get to the leaves of the constraint expression tree.
ItqlInterpreter
I had an attempt at a spike with some client side code, to see if I could build queries which were constrained on a blank node. Unfortunately, the query to retrieve the blank node kept failing on me. After trying all sorts of variations I tried selecting everything from the model. This worked. So then I constrained on a single column, and sure enough, nothing worked.
After lots of permutations I've discovered that ItqlInterpreter can execute a query easily and correctly. However, the combination of ItqlInterpreter.buildQuery, ItqlInterpreter.getSession() and Session.query() always fails. I was unable to work out why (not in the time I had, anyway). It must work for me on the server side because my session is pre-defined for me.
So I can't test my idea of querying on a blank node. I'll just have to do lots of logging at the server.
Posted by
Paula
at
Monday, May 30, 2005
0
comments
Saturday, May 28, 2005
SMOC
Coding over the last couple of days has gone well. I've been using the standard XP cycle of incrementally adding a bit, and running it to see that it works as intended. At this stage "working" is really a matter of making sure that the logs contain the data they are supposed to contain.
It's all been going well, but I'm only part way through. As Andrae likes to say, it's just a Small Matter Of Coding (SMOC). Meaning that the design is done, but the rest of the work is going to take some time. Now that I have the classpath issues resolved, I'm getting through a few hundred lines of code per day, so I'm happy the current pace. I just hope I don't run into any other major snags.
One issue that I've hit while building the Query objects has been the difference between a krule:Variable and a URIReference. This was bothering me when I wrote the RDFS, and the parsing of it made me consider it again.
Every time I run into constraint element I can ask for it's type. In the case of a variable, I can then use the name to create a new variable. Easy. But when I get a krule:URIReference I need to look for a krule:refersTo property to find the URI to construct the object. I hate if() constructs in code if I can avoid it. :-)
On the other hand, there is no real need for a variable to pick up a new attribute. If it did, then it should really be rdf:value.
This made me think about the krule:refersTo property that I've used for krule:URIReference, and so I looked up rdf:value again. I had thought that this was a datatype property (to use OWL parlance) instead of an object property, but the range is rdfs:Resource. Also compelling is the comment that the use of this property is encouraged to help define a common idiom. All of this was enough to make me change krule:refersTo to rdfs:Resource on krule:URIReference.
You'd think that would be enough to convince me to add a similar value to the variable, but I haven't. :-) The reason for this is usability and readability of the RDF for rules. Ideally I'd not have an extra indirection on the URI for URIReference either, but that has semantic consequences. I'd end up saying that some arbitrary URI has an rdf:type of krule:URIReference, which is just plain wrong.
The problem with having this difference between values and URI references is that each of them uses a different set of conjunctions to get all the data for construction. This means that I need two separate queries. To minimise the queries I'm doing, I'm initialising the whole rule reading procedure with a method that reads in everything of type krule:URIReference with their values, and mapping one to the other. This map is available later to any method that sees a krule:URIReference, and needs to get the URI it refers to.
Entailment and Consistency
Inferencing falls into two areas: entailment and consistency. I'm going to need to handle both. I haven't yet thought a lot about consistency checking, and that has worked out well.
Early on, I had thought that there would be a lot of consistency tests to perform, but after my learning exercise with cardinality I realise that consistency is not as strict as I'd thought. In the general case, if there is any possible interpretation that can make a data and its ontology correct, then it is consistent. This can rely on the most unlikely of statements which are not in the datastore. It is only when there are some direct conflicts that consistency becomes evident. Conflicting cardinality values, sameAs/differentFrom pairs, and the like are the most obvious examples. Fortunately, many of the less obvious examples will show up in the simpler tests after entailment has been performed, so the number of tests to be performed is not as daunting as it may first appear.
In the meantime, I've been learning more about logic at Guido's Wednesday sessions, which has also given me a better understanding of what I need to do.
As a first iteration, I'll be performing consistency checks only after entailment is complete. This will be more efficient, as it will only need to be done once, and will wait until all possible conflicting statements have been generated.
The problem with this approach is that it may be difficult to tell where the conflicting data came from. If entailed data conflicts with other entailed data, then the original data may be difficult to find (this is especially the case when the entailed data was entailed from other entailed data). The only way I can see around this is to perform checks after each entailment operation. This could get very expensive. Maybe it could be set as a debugging flag by the user?
Another thing to try during debugging will be more complicated consistency checks. Further entailments may show an inconsistency with simple tests, but it would be ideal if inconsistencies could be found and acted upon immediately. However, I believe that this approach could be of arbitrarily complexity, and only bound by the amount of work that the developer wishes to put into it. As a result, I don't think I'll be pursuing this for the time being. It would only be of use for debugging anyway, so the case to implement more complex test would have to be very strong (and involve a lot of $$$). :-)
Even with simple tests, I still need to define the consistency tests. The current rules are for entailment only, so I need to expand the vocabulary to handle them. The structure will be very similar to the entailment rules, so it should be easy to extend my current system to handle both. The main difference is when to run the rules, and to test for the existence of any result rows, rather than inserting results back into the model.
Posted by
Paula
at
Saturday, May 28, 2005
0
comments
Thursday, May 26, 2005
Unmarshalling
After integrating the rules interfaces I thought I'd run up the Kowari server and see that everything was still OK before proceeding. It wasn't of course (otherwise I wouldn't be talking about it here).
My error message was really helpful: Couldn't answer query.
So I started by grepping for this error message, and found it in several places within ItqlInterpreterSession. This was frustrating, as this means that I was seeing the problem at the entry point to the query process, rather than wherever the error was really occurring.
Of more concern was that each time this message was printed, the cause of the exception was supposed to be printed as well, but I was seeing nothing. I decided to print the message from the exception as well (not just the cause of the exception), only I saw nothing here either. Finally, in frustration, I changed each occurrence of the "Couldn't answer query" message to something unique, so I could tell which one was being printed. At this point the cause of the exception suddenly started to be printed as well.
I must have been running into a dependency problem where it wasn't building the new code modifications. I don't know why it suddenly started to work. If the new strings I'd created were not printed then I'd have known that it was a build script dependency problem, and I'd have done a clean build (I was on the verge of this already).
Given this problem I realised that I needed to perform clean builds more often in this debugging process. Incremental Kowari builds are already very time consuming, so performing a clean build each time meant that the rest of the debugging operation was guaranteed to be very time consuming (my latest clean build took 3 minutes, 17 seconds).
Anyway, I now knew that my problem was an "Unmarshalling" exception. So the problem was in RMI.
My first suspicion was that I'd updated a versioned interface at one end and not the other. The only changed interface was Session so I ran serialver to get the new number (the number generated by this program is essentially a checksum of relevant parts of the interface signature). However, the serial ID was exactly the same as before.
To confirm the problem I was seeing I commented out the new methods (buildRules() and runRules()), and tried to run another query. This took a while, as I had to make sure I got every implementation of the Session interface. Sure enough, it worked. So my problem was definitely in these methods, but where?
Wondering about the other rules classes I worked out which ones are to be transferred across RMI, and made sure that each was serializable (the exceptions were already serializable, but Rules was not). I couldn't see that this would make a difference, as I hadn't yet called a method that would move these objects, but I was trying everything I could think of. While I was at it, I made sure that each of these classes had serial IDs.
Once I had these changes throughout all the Session implementations I re-ran the code, and had no change in the output. I guess this was to be expected, but I had still be hopeful. :-)
While going through the various classes, it occurred to me that I had not yet made the necessary changes to RemoteSession. I wasn't yet passing the new function calls across RMI, so I thought that these would be safe to leave unimplemented for the moment. To intercept any inadvertent calls when the RMI interface was not complete, the RemoteSessionWrapperSession (blame Simon for the name, though it does make sense) was just throwing an UnsupportedOperationException for the new rules methods. With few other options to try, I decided to add the rules methods to the RemoteSession interface as well (after all, they were about to be needed anyway). Just a few extra lines, and they were implemented on SessionWrapperRemoteSession as well.
Unexpectedly, this did the trick! I still had an exception, but suddenly I had messages going over RMI, and the exceptions were being printed. With this information I discovered that the rule classes were not present in the classpath at the server end. Obviously they'd made it into the classpath during compilation, but when the distribution jar was built they were not being included.
Once rules-base-1.1.0.jar and krule-base-1.1.0.jar were included it all ran fine. I'm now back into the integration of rules with a database session.
Posted by
Paula
at
Thursday, May 26, 2005
0
comments
Monday, May 23, 2005
Whack-a-mole
Today was an exercise in frustration, but it was ultimately successful.
Every time I removed one dependency problem, another would raise its head. Some of the modules in Kowari have dependencies I'd never considered, and it was amazing just how many of these I had to deal with. Finally, I seemed to have them all under control, only to discover a new problem.
I've been reading data out of the kowari-config.xml file for some time, but I'd never tried to add to it. To start with, I naïvely thought that I could just add a new tag, and the code would automatically see it. This is not the case. I believe that Castor (the site is down ATM, but Google cached it just yesterday, so it can't be far away) is capable of just reading from an XML file and building an interface to what it finds, but that is not the configuration being used in Kowari. Instead, it uses an XSD file to describe the structure, and it builds the interface from that. At one point it was a DTD file and that is still to be found, but that is no longer used, and has fallen well behind.
Thanks to Simon for helping me find this schema.
I still had problems though, as my changes did not seem to be working. I added the new RuleLoader element, but nothing came up in the automatically generated Castor file. After a while I discovered that the field had resulted in it's own class files (RuleLoader.java and RuleLoaderDescriptor.java), but the field was still not showing up in the TucanaConfig class. Finally, I worked out what was missing in the XSD file. The TucanaConfig element wraps a list of references to other elements. All I had to do was enter a reference to the RuleLoader element, and it finally appeared in the TucanaConfig class.
Object Passing
One other thing that bothered me was when I tried to pass a session into the run() method of the rules engine. The rules interface can't know about sessions (because of dependency issues), so I thought I'd take the easy way out and pass it as an Object. However, for some reason the compiler refused to let me do this.
It's late, so maybe I missed something, but I got around this by providing a simple wrapper class for passing parameters to a rule framework for running. After all, different framework types might have difference requirements, so a generic way of passing parameters is probably a good thing. :-)
Posted by
Paula
at
Monday, May 23, 2005
0
comments
Sunday, May 22, 2005
Dependency Woes
Having had some success with the rules code sitting outside of Kowari (using RMI), I spent the weekend trying to bring it inside the project. Without that it was never going to scale.
Unfortunately, the rules code has led to a dependency nightmare. I need ItqlInterpreter to see an "apply" command, and so build a new rule framework based on the specified model. So the rules code needs access to a Session object in order to read the model. I then need to tell the database session to apply the rules (which in turn will be reading and writing on the session). So I suddenly have a dependency loop between rules and sessions. It is tempting to just bundle them together, but that doesn't quite do it either. As I mentioned the other day, I'm now using an ItqlInterpreter to parse my queries for me, so this leads to a loop between the iTQL module and the Kowari rules module.
The first step to address this has been to define a set of rule interfaces, and pull them into their own module. Then I defined a configurable rule engine item in the Kowari configuration file, to be loaded up via reflection, just like the resolvers and databases are. This solves some of the problems, and it also provides a perfect slot for a RuleML implementation at a later date.
I spent the entire weekend on this, and I'm still at it. Hopefully I'll be done in a few hours, as I really don't want to be bogged down in this any longer.
Posted by
Paula
at
Sunday, May 22, 2005
0
comments
Friday, May 20, 2005
Documentation
With so many changes coming in recently, I really need to document how they all work. I'm sure than Andrae will be encountering the same problems. However, it's not a trivial thing to do.
When Kowari was still managed by Tucana, Grant was formally writing everything up for us. We'd do the copy, while he would format it, structure the content where appropriate, correct our language, etc, etc. Every so often his files would get run through some proprietary program, producing the HTML files seen on the Kowari web site today.
None of this process is available now.
That leaves me with two choices. First, I can just edit the web pages in place. This is always a bad idea with automatically generated files, but if we're not generating them automatically anymore, then who cares? However, I'd like to think that we WILL be able to get this process going again eventually, and any manual changes to these files will be painful to deal with.
The second option is to find the program that Grant used, and generate everything again. The problems here are:
- Cost of the program (I have no idea how much it's worth).
- It's not open source.
- A bottleneck around whomever gets the job of building the documentation.
- Steep learning curve for whoever takes on the job.
- Will need a lot of work to make everything look consistent, getting publications out when needed, checking into Sourceforge, etc.
Either way, something has to be done. Does anyone have any ideas?
ItqlInterpreter
Who had the great idea of integrating the language parser with the execution framework? Let me explain.
When a client wants to execute a query, an
ItqlInterpreter object is used. This object parses the iTQL code, and sends the results to a database session for execution. The problem is that it integrates these two steps.I've written a lot of classes for representing rules when they are brought in from the data store, but the queries are causing me problems. My code is executing inside the database, so it will never need to go through Java RMI (establishing RMI connections is a bit bottleneck for Kowari). This means that I have a local
DatabaseSession that I want to use. Unfortunately, ItqlInterpreter uses an internal SessionFactory to get the session that it is going to use, meaning that it will always go across RMI.ItqlInterpreter creates Query objects, and I'd like to just use them on the Session that I have. This is why the coupling of interpreting and execution is such a problem for me.The easiest way around this would be to allow a
SessionFactory to be supplied to ItqlInterpreter, and to hand out the given DatabaseSession from it. However, I've been loathe to make this change without input from other people. I haven't seen much of them lately, so I have steered clear from this approach.Instead, I've been building up
Query objects manually, in the same way that ItqlInterpreter does it. There are numerous problems with this approach:- Takes a lot of work to write simple queries.
- The code is hard to write and read, and is therefore not extensible nor flexible.
- Difficult code like this is bug prone (OK so far, but I'm debugging a lot).
- It takes longer to write
ItqlInterpreter, and it's very fast.However, given my pace of progress, I'm thinking that I should just rip into
ItqlInterpreter and give it a new session factory, even without consulting the other developers.Top Down
I've started to understand the difficulties in representing hierarchal systems in databases. When building the object structure in memory, it is necessary to build objects from the top, and fill in their properties going down the hierarchy.
It is possible to come bottom up, but this can get arbitrarily complex, particularly when children at the bottom can be of arbitrary types. For instance, if A is a parent of B and C, then to build the object structure from the bottom-up, B and C must be found (actually, all leaf nodes must be found), and then you have to recognise that they share the same parent, and provide both of them when constructing A. This sounds easy enough, but gets awkward in practice, particularly when different branches have different depths.
So it is much easy to build this stuff bottom up. The problem there is that most of the classes being built expect to have all of their properties pre-built for them before they can be constructed. So that doesn't quite work either.
I'm dealing with it by building a network of objects which represents the structure, and which can be easily converted into the final class instances. It's relatively easy to build this way, and easy to follow, but again, it's more code than I wanted to write. Oh well, I guess that's what this job is all about. :-)
Posted by
Paula
at
Friday, May 20, 2005
2
comments
Saturday, May 14, 2005
Work, work, work
This has been a week of resolvers, testing, rules engines, and Carbon on Java.
Where to start? This is why I should be writing this stuff as it happens, rather than as a retrospective, but better late than never.
Prefix Resolver
I backed out the changes that I made to the string pool for prefix searching, but I've kept a copy. It's all been replaced with simpler code that uses Character.MAX_VALUE appended to the prefix string. This seems to work well with the findStringPoolRange() method from the string pool implementations.
I got this working with the prefixes defined in string literals, which makes conceptual sense, but it has some practical problems. In order to find all elements of a sequence, it is necessary to match a prefix of http://www.w3.org/1999/02/22-rdf-syntax-ns#_. When represented as a literal, this has to appear in its full form. However, if I use a URI instead, then I can take advantage of aliases, something that can't (and shouldn't) be done with string literals. This means that the prefix can be shown instead as <rdf:_>.
While the syntax for this is a lot more practical, it does cheat the semantics a little. After all, there is no URI of <rdf:_>, while there is a URI of <rdf:_1>. But the abbreviation is so much nicer that I'm sticking to it. I'm now supporting both URI and Literal prefixes, so anyone with a problem can continue to use the full expanded form as a string literal.
Backing the old code out, and putting in the dual URI/Literal option took longer than expected (as these things always do), so a lot of the week went in this, and the testing. However, it was worth it, as I can now select on a sequence. To do this I need a prefix model for the resolver (I'll also include the appropriate aliases here for convenience):
alias <http://tucana.org/tucana#> as tucana;
alias <http://kowari.org/owl/krule/#> as krule;
alias <http://www.w3.org/1999/02/22-rdf-syntax-ns#> as rdf;
alias <http://www.w3.org/2000/01/rdf-schema#> as rdfs;
alias <http://www.w3.org/2002/07/owl#> as owl;
create <rmi://localhost/server1#prefix> <tucana:PrefixModel>;The Kowari Rules schema uses a sequence of variables for a query, so I can use the prefix resolver along with the new <tucana:prefix> predicate to find these for me:select $query $p $o from <rmi://localhost/server1#rules>
where $query <krule:selectionVariables> $seq
and $seq <rdf:type> <rdf:Seq>
and $seq $p $o
and $p <tucana:prefix> <rdf:_> in <rmi://localhost/server1#prefix>;This has proven to be really handy for reading the RDF for the rules (which is why I wrote this resolver first). However, the real important of this is that I now have the tools to do full RDFS entailment. The final RDFS rule can now be encoded as:insert
select $id <rdf:type> <rdfs:ContainerMembershipProperty>
from <rmi://localhost/server1#model>
where $x $id $y
and $id <tucana:prefix> <rdf:_> in <rmi://localhost/server1#prefix>
into <rmi://localhost/server1#model>;So all I need now is the rule engine to do it. I spent the latter part of this week on just that. It looks like it's the parsing of the rules that takes most of the work, as the engine itself looks relatively straightforward. I'll see how this comes together in the next week.Meanwhile, I wrote a series of tests for the prefix matching (which I've mentioned several times in the last few weeks), and after many iterations of correcting both tests and code, I have everything checked in. Whew.
Namespaces
You'll note that the "magic" predicates and model types are all still using the "tucana" namespace. We really need to change that to "kowari". I'm a little hesitant to simply make wholesale changes here, as it will break other code. On the other hand, it is probably a good idea to do it sooner rather than later.
Does anyone have any objections to this change?
Warnings
Checking anything into Sourceforge still needs the full test suite to be run first (this is even more important now that we rarely see each other in person). During the tests I noted that there are some warnings about calls to
walk and trans. In each case they are complaining about variables rather than fixed values for a resource.In the case of
walk, then this will need to be addressed by allowing a variable, so long as it has been bound to a single value. This will be necessary in order to query for the node where the walking will start.For
trans then a bound variable of any type will need to be supported. This will make it possible to select all transitive statements in OWL.However, I'm aiming to have the basic rule engine with RDFS completed within the month. While these changes are vital, can I justify doing them now? They are not needed for RDFS, so perhaps I should postpone it. I'm just reluctant to put it off, as the change is important and should only take a few days to do.
Maybe I should put my other projects on the back burner and start using my "after hours" time to get this done instead? That'd be a shame, as I need some sanity in my life.
jCarbonMetadata
As my code for wrapping the Carbon classes became more complete and refined, I decided to put it up on SourceForge. Since it was just supposed to wrap the Carbon metadata classes for Java, I called it jCarbonMetadata. However, I'm starting to wonder if the scope is shifting.
The latest part of this project has been to implement the query class. To properly write the javadoc for this code, I really need to duplicate Apple's documentation, but I don't think I'd be allowed to do that. However, I'm not sure about this. I'm providing access to Apple's classes, so it makes sense to use Apple's docs. Writing my own descriptions sounds like a recipe for disaster, particularly as I'm often doing this stuff after midnight. Linking isn't very practical either.
I've been making progress anyway, writing wrapper functions around many of the
MDQuery functions. Just when I thought I was near completion I discovered a method that threw my whole plan into chaos.The initial idea was to write a wrapper around the three relevant MD classes:
MDItem, MDQuery, MDSchema. The methods for these classes all return strings and collection objects from the CoreFoundation framework. Rather than provide a wrapper implementation for each of these objects, I've been converting everything to its java equivalent and returning that instead. Not only does this mean less work (as I have fewer classes to implement), but it makes more sense for a Java programmer anyway, as the Java classes are the ones they need to be using.This was all going well until I got to the function called
MDQueryCopyValuesOfAttributes. This function returns a CFArray of values. I was expecting to convert this into a java.lang.Object[] until I read the following:The array contents may change over time if the query is configured for live-updates."
So if I just copy the results into an array I'll lose this live-update functionality!
There are two approaches I can take here. The first is to just have a static interface, and wear the loss in functionality. The
CFArray has to be polled for a change in size anyway, so why not poll the function call instead?The second approach is to wrap a
CFArray in another Java class. While this would take more work, it is a more complete solution.For the moment I'm hedging my bets, and have written two methods. The first method will return an array, while the second will return a dynamically updated
MDQuery object. I have a stub class for this at the moment, though I don't expect to flesh it out until everything else is done. I thought that the method that returns the Object[] might just use the CFArray method, and have this class provide a toArray() method, but in the end I decided that I can cross the JNI boundary less often if I provide separate methods in C.If I'm going to look at
CFArray I figured that I might as well find out what else it offers. Most of the functions are as you'd expect, but there is also a function called CFArrayApplyFunction for applying a function to each element of the array (this brings flashbacks to the STL for me). That seems relatively useful, so I've starting thinking about the rest of the Core Foundation (CF).Apple have explicitly said that Carbon is not going away, and indeed that there are some things that can only be done using Carbon. With this in mind, maybe there is room to implement the whole of CF in Java? I don't think any of it would be too hard, though it would take a lot of time.
This made me wonder if I should really be converting all of the CF objects into their Java equivalents. Perhaps I should implement each of these objects in Java, and return those instead. Then they can be converted into the Java classes by appropriate pure Java functions.
Considering this a little more, I decided that even if I implement all of these classes, then I'd still do conversions in C, through a JNI interface. There are two reasons for this. The first is that some objects must be converted in C, such as numbers, and for this reason most of the conversion functions have already been written. The second reason is to minimise crossing the JNI boundary. Converting objects like
CFArray and CFDictionary in Java would mean repeatedly calling the same JNI access functions.As I said earlier, I'll just stick to the three main classes, and consider the remaining CF classes when I get there. I still have other projects on the go, so I'll just have to consider the importance of these things once I get there. For the moment I just want this code for a Kowari resolver.
I'm sure there's other stuff to talk about, but I'd better get back to work.
Posted by
Paula
at
Saturday, May 14, 2005
0
comments
Sunday, May 08, 2005
Redundant Code
Last night at the reception I was talking to DavidM about my JNI implementation for OSX metadata. David thought that this was unnecessary, as Apple have a full set of Cocoa bindings for Java. If this were true, then the code I wrote would be redundant. While it was great practice, I'd rather that I'd provided something really useful.
BTW, the wedding went really well. No honeymoon yet, but maybe we'll try to organise something later in the year.
Tonight I started looking for documentation on Java-Cocoa, to see if David is right. It turns out that he is (mostly). Damn. :-)
Looking for documentation on Apple's site tells me that Java-Cocoa does exist, but it seemed hard to find. Looking in the Xcode documentation, I found it under:
ADC Home > Reference Library > Documentation > Cocoa > Java
Google searching also finds references to it (often in mailing lists), but typically in disparaging tones. The complaints have been that the code is buggy, and that documentation is poor.
So I went looking for it on the local system. The first place I looked was in /System/Library/Java. The Extensions subdirectory looked promising, but wasn't. However, I soon found all of the "NS" Cocoa classes in com/apple/cocoa/foundation. This included NSMetadataItem, which I thought was the exact equivalent of what I'd just written.
Oh well. I'm still glad for the coding practice. :-}
However, looking at the situation more carefully, it's not so clear cut. It seems that the NSMetadataItem has no constructor that accepts a file. Instead, these objects are instantiated as the result of a metadata query. So from what I can see, the MDItem class from Carbon allows the metadata from a given file to be retrieved, but there is no equivalent functionality in Cocoa. It may exist, but I haven't found it yet.
Once I knew what I was looking for, I started searching for the Java API for this class, but with no luck. Most other classes are available, but none of the metadata classes. Finally, I tried using the javap disassembler on the NSMetadataItem, getting this:
Compiled from "NSMetadataItem.java"
public class com.apple.cocoa.foundation.NSMetadataItem extends com.apple.cocoa.foundation.NSObject{
public native java.lang.Object valueForAttribute(java.lang.String);
public native com.apple.cocoa.foundation.NSDictionary valuesForAttributes(com.apple.cocoa.foundation.NSArray);
public native com.apple.cocoa.foundation.NSArray attributes();
protected com.apple.cocoa.foundation.NSMetadataItem(boolean, int);
public com.apple.cocoa.foundation.NSMetadataItem();
static {};
}This confirms that there is no constructor which accepts a file path.I also noted that these classes return
NSDictionary and NSArray objects. This makes the wrapping a little thinner than it need be, as Java code would usually use java.util.Map and Object[] for the same purposes. At least it uses Java strings, rather than NSStringReference.So I'm thinking I'll keep up with this code for two reasons. The first is the more convenient interfaces from
java.util. The second is that having a Carbon wrapper seems to offer a couple of functions that are unavailable under Cocoa (like getting metadata from a specific file).This extra functionality in Carbon seems to be confirmed by looking at the contents of the "mdls" command (using "nm"). The symbols it uses all indicate that the code is written in, and linked to, Objective C. However, it still uses
MDItem from Core Framework. There shouldn't be a need for this, but with no support for file-specific metadata in Cocoa then it becomes essential.I'll keep writing with this code for the time being. Given the lack of documentation, I suppose that Apple are still in the process of working on this, and will probably release something more complete as 10.4 progresses. In the meantime, I find this sort of thing fun, and everyone needs a hobby. :-)
Posted by
Paula
at
Sunday, May 08, 2005
2
comments
Friday, May 06, 2005
MDItem
Well it's the night before the wedding, and I should be trying to make sure everything is ready. Instead I've finished the first JNI wrapper class for the OS X.4 metadata classes. It hasn't been easy finding the time, as I was trying to work on real Kowari stuff, as well as trying to help Anne with all those last minute details.
The class I've finished is MDItem. This is a direct rip-off of the MDItem class found in the Core Framework. There is a one-to-one correspondence between the classes' methods. There are a couple of minor TODO items, and I should add in a list of the common attributes, but it's basically done.
A lot of work went into the code that converts Core Framework objects into Java objects, and back again. This code will get re-used heavily, so any further work will be a lot faster.
Speaking of further work, the main class I need to work on now is MDQuery. I expect to start on that next week. After I have that one I should be able to start on the resolver! There's also MDSchema, but I don't think it's as important to get things running. I also can't see the need to try and write a file importer in Java, as the file system is not going to want to start up a JVM. :-)
Anyway, if anyone has OS X.4 and is interested in trying it out, then you can find the Xcode project here. The class comes with a main() to test it out, so you can run the resulting jar like this:
$ cd buildThe output is in two parts. The first part is all of the attributes, and gives the same results as the
$ export DYLD_LIBRARY_PATH=${DYLD_LIBRARY_PATH}:.
$ java -jar metadata.jar libmetadata.jnilib
mdls command. The second part is a repeat of the first part, but only tries to pick up every second attribute instead (this is to test a different method). You'll note that I ran it on the library itself, but you can run it on any file in the system.I hope someone likes it. :-)
Posted by
Paula
at
Friday, May 06, 2005
0
comments
Thursday, May 05, 2005
Time
It's still hard to find time to write at the moment. Anne and I are getting married tomorrow, so hopefully I'll have a little more time next week.
In the meantime, we've had several public holidays recently (we get them all at one time of the year). Unfortunately I've only taken advantage of the most recent of them, which was last Monday.
Anyway, with the wedding tomorrow and a lot of work to get done, I won't be writing much today either. I just thought I'd better write something before I get too far out of the habit.
Uni
I did my confirmation last Friday. The internal web page with the administrative details of confirmation has been taken down, so it wasn't until the last moment that I discovered when the paper was due. It was sooner than I had thought, so I didn't have a lot of time for feedback on it. As a result, no one got back to me in time. It was frustrating, but at least the paper was accepted.
Again, with lack of time, I didn't get a lot of time to prepare for the presentation either. I'm normally good with presentations, so I was feeling a bit overconfident about it. But once I was in there I realised how out of practice and under prepared I was. Fortunately, confirmation is evaluated in relation to the quality expected of most students, so from that perspective I did well. All that matters is whether I was accepted or not, and I did that easily.
However, I was very disappointed with myself, as I can do a lot better. I've asked Bob for the opportunity to give several more presentations, to get some more practice.
Resolvers
I backed out some of the comparator changes, as I mentioned last time. However, I've kept them archived, as I still feel that they are more "correct". If I ever run into problems with the current method I'll have some code to fall back onto. This code is quite long, so I wouldn't want to write it again.
Debugging has gone slowly, with everything else I've been doing, but it's all looking good. The main problem I seem to have now is a need for a resolver that will also search along collection lists. I have to think about the consequences of this though, as this would remove duplicate entries. None of the use cases I have contain duplicates, but it would still be a problem in some circumstances. Maybe I should just document this as the required behaviour (which it is) and move on.
In my spare time (like I have any of that) I've also started work on another resolver. This time it's not related to my project, but I think it's going to be important...
OSX
Mac OS X.4 was supposed to arrive last Friday, but it didn't. If I'd realised it might be late (Apple guaranteed delivery on the release day) I wouldn't have pre-paid for it, and would have gone off to an Apple Centre to buy it instead. This is a toy I've been looking forward to. :-)
At least this gave me a chance to play with, and set up my new Neuston MC-500. Linux does not synchronize it's video output too well, so fast-action video scenes can tear a little. Also, MythTV hates sending output to the TV through my old MGA G400 card (I don't play games, so I haven't had a need to upgrade it). So now I have MythTV doing the recording, and the Neuston doing the playback. The user interface with the remote is easy for Anne to use, so she's been able to show Luc a lot of recorded "Wiggles". :-)
But I digress...
The real reason I've been waiting for OSX has been to use Spotlight. This is a hash-based index for meta-data of all the files on disk. For some time I've been thinking of implementing something like this for Linux or OSX, but now that Apple has done it for me I don't need to. :-) Longhorn will be doing something similar as well, so it should not be long before this becomes a common feature for modern file systems.
I was thinking about using Kowari (or a similar system written in C) to implement the filesystem index. Now that I don't have to, I've been thinking of inverting the whole thing, and using the filesystem index as a backend for Kowari. This just requires a resolver that wraps Apple's Spotlight interfaces. That's the resolver I've been writing.
For the moment, I've been using JNI to implement a set of Java classes which provide the meta-data interfaces from Apple's "Core Foundation". This has been a good way to learn Core Foundation. It's also good practice in JNI, as I'm trying to do everything "right". This means a lot of error checking, etc.
Eventually I expect to provide a set of metadata classes for Longhorn as well. However, I should see what functionality is provided there before I try to work out a common interface that will work across filesystems. In the long run, I'd like to have library that detects the OS, and automatically loads up the correct JNI library. I could even have a fallback class written in Java that returns basic meta data about files (filename, timestamp, owner, etc).
But at this stage I just have classes which work on Spotlight. Once I've wrapped them in a resolver I'll be adding them to Kowari.
If anyone is interested in this code, then just let me know. It's all being open-sourced, but I just haven't got around to publishing it yet. I'm been doing this Spotlight work in my own time (since I'm working on the rules engine in the day), and between the confirmation, the wedding, and last weekend's triathlon, I haven't had a lot of "spare time" at all. :-)
Posted by
Paula
at
Thursday, May 05, 2005
0
comments