Jskad

Author	SHA1	Message	Date
dchandler	7198f23361	I really hesitate to commit this because I'm not sure what it brings to the table exactly and I fear that it makes the ACIP->Tibetan converter code a lot uglier. The TODO(DLC)[EWTS->Tibetan] comments littered throughout are part of the ugliness; they point to the ugliness. If each were addressed, cleanliness could perhaps be achieved. I've largely forgotten exactly what this change does, but it attempts to improve EWTS->Tibetan conversion. The lexer is probably really, really primitive. I concentrate here on converting a single tsheg bar rather than a whole document. Eclipse was used during part of my journey here and some imports were reorganized merely because I could. :) (Eclipse was needed when the usual ant build failed to run a new test EWTSTest. And I wanted its debugger.) Next steps: end-to-end EWTS tests should bring many problems to light. Fix those. Triage all the TODO comments. I don't know that I'll ever really trust the implementation. The tests are valuable, though. A clean implementation of EWTS->Tibetan in Jython might hold enough interest for me; I'd like to learn Python.	2005-06-20 06:18:00 +00:00
dchandler	c16f633ecf	Two things: One, TMW->EWTS gives dbas and dngas instead of dabs and dangs because Chris Fynn's e-mail from today has dbas and dngas. Second, Down with ACIPRules. Long live ACIPTraits. EWTS->Tibetan conversion is closer still.	2005-02-22 04:36:54 +00:00
dchandler	4c268c5ea2	Refactored so that there can be an EWTS scanner and an ACIP scanner.	2005-02-21 05:37:01 +00:00
dchandler	3e0168b384	Renamed ACIPConverter to TConverter. Added a needed parameter (the only needed parameter in that class's interface AFAIK.	2005-02-21 01:35:23 +00:00
dchandler	37bf9a736d	I did this stuff back in August. It's all in support of EWTS->Tibetan conversion. The tag 'TODO(DLC)[EWTS->Tibetan]' exists all over the place. EWTS->Tibetan isn't here yet; lexing isn't here yet; this is mainly a refactoring so that the ACIP->Tibetan code can be reused to do EWTS->Tibetan. I'm committing this because tests pass (it shouldn't be breaking anything), because I want a checkpoint, and because the laptop this sandbox was on isn't my preferred development environment.	2005-02-21 01:16:10 +00:00
dchandler	9025fb42d6	TMW->EWTS 998476 partial fix: "aM" is generated now correctly. Before you got "M".	2005-02-07 04:00:42 +00:00
dchandler	8dcb623382	TMW->EWTS: Fixed part of bug 998476 and part of an undocumented bug. Discovered a new bug, "aM" should be generated but only "M" is. The undocumented bug was that laMA was generated when lAM should have been. The part of bug 998476 that was fixed: laM, laH, etc. are now generated. This does nothing about paN etc. Some refactoring here; this is not a minimal diff. Added tests of TMW->EWTS that use ACIP to get the TMW in place because EWTS->TMW is a faulty keyboard at present.	2005-02-07 03:17:40 +00:00
dchandler	7acbce3361	Added errors 142 and 143, which are produced when converting yig chung to a Unicode text file, which cannot support font size changes.	2004-06-06 21:59:16 +00:00
dchandler	df262aa148	It is now a compile-time option whether to treat []- and {}-bracketed sequences as text to be passed through (without the brackets in the case of {}) literally, which is the case by default because Robert Chilton requested it, or the old, ad-hoc mechanism which could be useful for finding some ugly input. Made a couple of error messages a little more verbose now that we have short-message mode.	2004-06-06 21:39:06 +00:00
dchandler	8a9271a3d8	I broke warning 507 into two warnings, one high-priority (512) and one low-priority (507).	2004-05-01 20:49:53 +00:00
dchandler	1a055f3472	I don't think warning level "None" was really doing the trick. Fixed that. You can now customize the severities of all warnings, even 504 and 510. When warning level is "None", scanning, i.e. lexical analysis, is faster.	2004-04-25 00:37:57 +00:00
dchandler	e2d42f36eb	Robert Chilton's experience inspired me to make the handling of errors and warnings in ACIP->Tibetan conversion much more configurable. You can now choose from short or long error messages, for one thing. You can change the severity of almost all warnings. Each error and warning has an error code. Errors and warnings are better tested. The converter GUI has a new checkbox for short messages; the converter CLI has a new mandatory option for short messages. I also fixed a bug whereby certain errors were not being appended to the 'errors' StringBuffer.	2004-04-24 17:49:16 +00:00
dchandler	0ee90a0fb0	Added many ACIP->TMW->ACIP tests. They found no bugs.	2004-04-17 17:28:26 +00:00
dchandler	adcf9de952	Two new tests.	2004-04-17 15:14:46 +00:00
dchandler	7eca276a62	TMW->Unicode conversions have changed; now using U+0F6A for the stacks whose EWTS transliteration begins with "R+". ACIP->* conversions and test baselines were updated to deal with the "r+..."=>"R+..." change.	2004-04-10 16:03:25 +00:00
dchandler	76356f4009	ACIP->Tibetan now gives an error when {?} is seen alone (not in {[?]} or {[*FOO?]}, but alone). Bug 860192 is fixed.	2004-03-15 00:49:01 +00:00
dchandler	c1aa81e943	RFE 860190: ACIP->Unicode now gives a warning when it outputs something that can't be represented in TMW.	2003-12-16 07:45:40 +00:00
dchandler	848349fd3a	More tests.	2003-12-15 08:16:06 +00:00
dchandler	e7a9e7968f	ACIP->Unicode now uses two characters for consonants instead of one. This matches the dislike for characters like U+0F77 etc. ACIP->Tibetan was not giving an error for BCWA because it parsed like BCVA. Fixed.	2003-12-15 07:32:14 +00:00
dchandler	8664571577	Warnings were not being detected correctly. Fixed. ACIP->Unicode uses U+0020, ' ', for whitespace. ACIP->TMW uses the TMW whitespace for whitespace.	2003-12-14 08:38:10 +00:00
dchandler	76c2e969ac	Fixed ACIP->Unicode bug for YYE etc., things with full-formed subjoined consonants and vowels. Fixed ACIP->TMW for YYA etc., things with full-formed subjoined consonants.	2003-12-14 07:36:21 +00:00
dchandler	581643cf59	{DAN,\nLHAG} used to be treated like {DAN, LHAG} but that got broken. Fixed. Added tests for lexer's handling of ACIP spaces etc.	2003-12-10 06:55:16 +00:00
dchandler	a466bad939	ACIP->TMW now supports EWTS PUA {\uF021}-style escapes. Our extended ACIP is thus TMW-complete and useful for testing.	2003-12-08 07:51:45 +00:00
dchandler	c43e9a446b	Revamped some ACIP->Tibetan error messages.	2003-12-06 20:19:40 +00:00
dchandler	c9c771d1ee	ACIP {&}, as in {KO&HAm,}, is supported.	2003-11-30 02:18:59 +00:00
dchandler	ac412c994b	Now {Pm} is treated like {PAm}; {Pm:} is like {PAm:}; {P:} is like {PA:}.	2003-11-30 02:06:48 +00:00
dchandler	4e6a9c299f	ACIP % {MTHAR%} and o {Ko} and ^ {^GONG SA} are now supported. A % always causes a warning.	2003-11-11 03:43:11 +00:00
dchandler	04816acb74	ACIP->Unicode was broken for KshR, ndRY, ndY, YY, and RY -- those stacks that use full-form subjoined RA and YA consonants. ACIP {RVA} was converting to the wrong things. The TMW for {RVA} was converting to the wrong ACIP. Checked all the 'DLC' tags in the ttt (ACIP->Tibetan) package.	2003-11-09 01:07:45 +00:00
dchandler	94a43d3f39	Now anything not clearly native Tibetan is colored green when coloring is enabled. G'EEm is "native", though -- the only "vowel" that implies non-nativeness is {:}, as in {KA:}.	2003-10-26 18:56:48 +00:00
dchandler	e74547d743	GA-YOGS now parses like G-YOGS and GAYOGS do.	2003-10-26 18:06:38 +00:00
dchandler	31b3020d07	Added a test case that runs almost all the tsheg bars from all non-reference, publicly available ACIP files (hundreds of megabytes of them) through the converter. The frequencies of these tsheg bars in in the file, too.	2003-10-26 06:02:48 +00:00
dchandler	6bda550157	The ACIP "BNA" was converting to B-NA instead of B+NA, even though NA cannot take a BA prefix. This was because BNA was interpreted as root-suffix. In ACIP, BN is surely B+N unless N takes a B prefix, so root-suffix is out of the question. Now Jskad has two "Convert selected ACIP to Tibetan" conversions, one with and one without warnings, built in to Jskad proper (not the converter, that is).	2003-10-26 00:32:55 +00:00
dchandler	306cf2817c	Private correspondence with Robert Chilton led to me to add and remove a few prefix rules. BLC and BGL are here, BLK, BLG, BLNG, BLJ, BNG, BJ, BNY, BN, and BDZ are gone. Added a few new tests.	2003-10-25 21:47:34 +00:00
dchandler	7d24ab393f	Code cleanup.	2003-10-21 03:44:02 +00:00
dchandler	c764eee8d0	Added a new warning for DMAR and others affected similarly affected by prefix rules, where seeing D+MAR, not D-MAR, could have caused an input operator to type in DMAR. This is a "Most" warning, but DMA causes a higher-priority "Some" warning.	2003-10-21 03:36:57 +00:00
dchandler	2f39921381	Added more test cases.	2003-10-21 02:14:45 +00:00
dchandler	2f81a801ef	Added three new kinds of warnings to ACIP->Tibetan conversions.	2003-10-21 02:00:49 +00:00
dchandler	3aa3859354	ACIP->Unicode crash fixed. 5% of the code for support of ACIP->Unicode.rtf is here.	2003-10-19 22:19:16 +00:00
dchandler	5aab4acc93	I've undone the SNYAM'AM == SNYAMA'AM hack. The only occurrence of SNYAM'AM in the ACIP texts I've got is likely a typo, says Robert Chilton. The code would be cleaner if I could bear to delete my terrible hack. Maybe in a month, when I don't feel so dumb for coding it up in the first place. The correct solution for such things is to give the ACIP->Tibetan converters a pre-filter mechanism. This would be before the lexer or part of the lexer (maybe you only want to filter tsheg bars), and it would allow the end user to specify things like "s/SNYAM'AM/S+NYAMA'AMA/g".	2003-10-19 20:48:22 +00:00
dchandler	4b1395e0ba	Jskad has a new feature: Convert Selection from ACIP to Tibetan. It uses the ACIP converter to do its work. Improved some error messages from the ACIP->Tibetan converter.	2003-10-19 20:16:06 +00:00
dchandler	0edebd55d7	We were dying in the "can ts+h take a ga prefix?" check for GTZHAN.	2003-10-19 03:47:33 +00:00
dchandler	557ed7ed44	DKY'O etc. weren't being handled properly by ACIP->Tibetan. Now they are.	2003-10-18 17:49:29 +00:00
dchandler	3b55ea509f	Prefix rules have changed. A few are gone; a few new ones are here. I've implemented here a list that Robert Chilton sent me in private correspondence. He doesn't describe it as definitive, but since it affects ACIP->Tibetan conversions, and it's the best I've got, here they are. There's still an optional warning about "Hey, prefix rules matter for this tsheg bar." I've left in a few rules that I didn't find on RC's list; I've asked him to look into these further.	2003-10-18 05:48:53 +00:00
dchandler	8c99adeb63	TMW->EWTS, TMW->ACIP, and ACIP->Unicode/TMW now support more appendages. Personal correspondence with Robert Chilton led me to support, besides 'am, 'ang, 'o, 'i, and 'u, the following: 'e (used in foreign transliteration) 'ongs 'is 'os 'ur 'us 'ung	2003-10-18 03:04:47 +00:00
dchandler	5e18feb47d	ACIP now stacks greedily. TTTTTA is T+T+T+T+TA, even though that stack doesn't exist in TM or TMW. Robert Chilton, in personal correspondence, agreed that this is the way to do things. ACIP handles the appendages 'AM, 'ANG, 'US, 'UR, 'I, 'O, and 'U correctly.	2003-10-16 04:15:10 +00:00
dchandler	115d0e0e6c	Fixed ACIP->TMW vowels like 'I etc. Fixed ACIP->Unicode/TMW for BDE, which should be B-DE, not B+DE, because the former is legal Tibetan. The ACIP->EWTS subroutine has improved. TMW->Wylie and TMW->ACIP are improved in error cases. TMW->ACIP has friendly embedded error messages now.	2003-09-12 05:06:37 +00:00
dchandler	16817d0b8e	Fixed Javadocs.	2003-09-10 01:19:05 +00:00
dchandler	d8657abd44	ACIP font shrinking as in {KA (GA)} is now supported.	2003-09-07 18:30:59 +00:00
dchandler	07e360d9a8	The ACIP {NYA%} is supported. {NYAo} and {NYAx} are confusing to me, because I don't know which glyphs o and x correspond to. For that reason, they cause ERRORs. The proposed THDL Extended Wylie ~X and X is now used for U+0F35 and U+0F37 respectively.	2003-09-07 16:19:50 +00:00
dchandler	cc615f34df	ACIP->TMW and ACIP->Unicode have my pre-stamp of non-approval. Except for (NYAx} and {NYAo}, they're as good as I'll get them without input from experts of the employ of a complementary, syllabary-based approach.	2003-09-04 04:34:18 +00:00

1 2

60 commits