kodirovka sooschenij.

ZXNet echo conference «zxnet.pc»

From Kirill Frolov To All 31 December 2003

Press RESET immediately, All! I wrote a letter to ZX.SPECTRUM, went to this conference and immediately saw a couple inappropriate letters. I decided to "forward". In short: “+7 FIDO”, “IBMPC” and similar heresies do not have encodings right to exist. You must use an encoding that is exactly there are Russian letters. For example, CP866. ASCII does not contain Russian letters. There is no such thing as high-ascii, ascii is by definition a 7-bit code. Newsgroups: fido7.zx.spectrum From: Kirill Frolov X-Comment-To: All Subject: Speaking Russian -- ZX Spectrum Date: Wed Dec 31 09:38:24 MSK 2003 Press RESET immediately, All! On Tue, 30 Dec 03 17:39:40 +0300, Andrei Denisenko wrote: AD> A long time ago, in the 20s of the 20th century, Slavik Tretiak was building a machine gun on AD> about: Р ї╜∙Г∙╛ ZX Spectrum Pojaluista po-russki. This is what non-compliance with certain technical requirements leads to. Let me quote both together with official information, on which and will focus your attention. The letter is addressed to ALL ZX.SPECTRUM SUBSCRIBERS. Please be careful and draw your own conclusions, check settings for your programs. So, the original letter (headers only): Newsgroups: fido7.zx.spectrum From: Slavik Tretiak X-Comment-To: Andrei Denisenko Subject: Development of ZX SpectrumDate: Tue, 30 Dec 03 11:41:02 +0300 Message-ID: <1072806062@p33.f2.n451.z2.FidoNet.ftn> X-FTN-MSGID: 2:451/2.33@FidoNet 3ff1b8ae X-FTN-CHRS: LATIN-1 2 X-FTN-Tearline: Spot 1.3b #118 This letter WAS NOT read by Andrey Denisenko because the software he used showed it in ISO-8859-1 encoding, while how the message text itself was encoded in CP866 encoding. The program did *absolutely correctly* and in accordance with the relevant standards (FTS) and recommendations (FTSC) of the Fidonet network. It's all about in the service information line ("kludge") "X-FTN-CHRS: LATIN-1 2", informing about the encoding used in the letter. (Note: fidonet users may see this string as "kludge" "@CHRS: LATIN-1 2", where the '@' character is a non-displayable character with code 01) This string "kludge" contains INCORRECT TECHNICAL INFORMATION, therefore that in reality the message is encoded in a different encoding. I would like to note that this letter could not contain messages in Russian, since the LATIN-1 encoding does not contain the necessary characters. Now let's look at Andrei Denisenko's letter, which was correct must have been read by all teleconference subscribers (only headers): Newsgroups: fido7.zx.spectrum From: Andrei Denisenko X-Comment-To: Slavik Tretiak Subject: Рї╜∙Г∙╛ ZX Spectrum Date: Tue, 30 Dec 03 17:39:40 +0300Message-ID: <4215340360@p11.f7.n5064.z2.ftn> X-FTN-MSGID: 2:5064/7.11 fb40fd48 What is visible here? The distorted theme immediately catches your eye message, the original was “Development of the ZX Spectrum”. This is explained the fact that Andrey's software is guided by something that does not correspond to reality technical information incorrectly recoded the text. In addition, it is clear that the CHRS “kludge” is missing. That is, it is unambiguous to determine message encoding is not possible. Such messages have every right to exist in the Fidonet network and meet all the necessary requirements standards. However, programs processing such messages must make any assumptions about the text encoding system. Usually it is assumed that the one generally accepted for a given network or region is used or network zone encoding. This works well for messages written in the language accepted in this part of the network, for locally distributed letters. But when sending mail (NETMAIL) to another zone, in international teleconferences, or in teleconferences where it is customary to use multiple languages, this may cause a problem -- assumptions letters about encoding may be false. So there is a problem: an undefined message encoding. And there is for its solution provided by Fidonet standards - use lines of technical information of the message header ("kludge") "CHRS". This "kludge" must contain the official nametext coding systems. For the Russian part of the Fidonet network It is accepted to use CP866 encoding. Ukraine, Belarus and other republics of the former USSR may use other encodings, only partially identical to CP866, for example CP1125. In European parts of the 2nd zone of the Fidonet network and in the 1st zone (USA) it is common to use CP855 or ISO-8859-1 encodings (may be referred to as LATIN-1). Theoretically, a letter can be encoded in any of the broad known (not necessarily standard - standard are only encodings with the ISO prefix, for example, for Russian language, only the ISO-8859-5 encoding is defined - this is the standard number), including, the message can be encoded in unicode using UTF-8 encoding system (use of multibyte encoding systems unicode, like USC-16 is not possible in Fidonet). And, again theoretically, messages encoded in any of the known encodings and containing correct "kludge" CHRS, can be displayed correctly on the screen recipient. So when using unicode it is possible to use The letter contains several languages, for example Japanese and Russian. But that's all in theory, in practice this is hampered by a number of restrictions imposed outdated software that does not support work with various encodings. This is why it is customary to encode messages in some encoding generally accepted for the network, so that all subscribersnetworks, including those using outdated software, could easily read your message. So, what conclusions can be drawn from what has been said: “kludge” CHRS MUST CONTAIN RELIABLE INFORMATION, or be missing at all. Otherwise, your letter may be misinterpreted. If this “kludge” is missing, the letter MUST BE ENCODED IN A COMMON ENCODING FOR YOUR NETWORK. If "kludge" present IT IS NOT ADVISABLE TO USE CODINGS OTHER THAN COMMONLY ACCEPTED IN YOUR NETWORK, because not all subscribers are ready to accept letters like this. So, in particular, a large network node 2:5020/400, through which internet users have the ability to access teleconferencing of the Fidonet network, uses software does not support text recoding according to the CHRS "kludge", and always expects that your letters will be encoded in generally accepted for network encoding. If you encode your letters in excellent from the generally accepted encoding, users of internet news servers will not be able to read your letters. REMEMBER: IF THE CHRS KLADGE CONTAINS INCORRECT INFORMATION, YOUR THE LETTER MAY BE INCORRECTLY DECODED BY THE MESSAGE RECIPIENT, SO IT IS DISTORTED AT INTERMEDIATE NODES_ WHICH CLEARLY FOLLOW STANDARDS. Recommendations from the author of this letter for setting up your software canbe like this: make sure that the CHRS "kludge" is present and always contained correct information. So for Russia and most of the former USSR this will be the generally accepted CP866 encoding. Try without extreme pressure the need not to write letters in an encoding different from the generally accepted one. If you cannot ensure that the CHRS "kludge" is correct, as in the case the first of the messages discussed in this letter, MANDATORY make sure that this “kludge” is not present at all in your messages (even if the software you use does not allow it, "kludge" can be cut with third-party programs), and always encode your letters into a generally accepted encoding. If you need to put text in language, which cannot be represented by a generally accepted encoding, encode it in UUE or Base-64. Many subscribers may have questions about setting up their programs. First of all, carefully study the documentation, if it no, then think about changing the software. So the latest version of the editor GoldED (1.1.4.7 or later) fully supports various encodings, but does not work with unicode. You can ask questions at Fidonet network teleconferences "SU.CHAINIK" or "RU.UNIX.FTN" (for using unix-like systems). As an attachment to the letter, I provide document FSP1013, which describes issues of internationalization in the Fidonet network. It might not be the last document and contains inaccuracies.********************************************************************** FTSC FIDONET TECHNICAL STANDARDS COMMITTEE ********************************************************************** Publication: FSP-1013 Revision: 1 Title: Character set definition in Fidonet messages Author: Peter Karlsson Revision Date: 04 September 1999 Expiry Date: 04 September 2199 Contents: 1. Introduction 2. Format of the identifier 3. Supported levels 4. Supported character sets 5. Obsolete identifiers 6. Notes Status of this document This document suggests a proposed protocol for the FidoNet(r) community, and requests discussion and suggestions for improvements. Distribution of this document is unlimited. Abstract This document defines the identifiers that are used to indicate the character set used within messages that are distributed in Fidonet. There has been many attempts on defining a common standard on what character set is used in Fidonet distributed messages. The only one that has gained widespread is the "CHRS" standard, described in FSC-0054. This document tries to describe the current usage, as well as standardizing the parts of it that was ambigously defined. 1. Introduction As Fidonet is an international network, one has to consider that not all people speak English. Many languages also have alphabets that are either bigger than the standard English alphabet, or completely different. To keep track of which character set is used for a particular message, this documents describes a way to identify this to the reader. 2. Format of identifier The CHRS control line is formatted as follows: ^ACHRS: Where is a character string of no more than eight (8) characters identifying the character set to use, and level is a numeric value describing what level of CHRS this message is written in. For backward compatibility, the use of "CHARSET" instead of "CHRS" should be supported when reading messages. Some current implementations do not add the required field. Implemetations of this document should allow for this usage, but never write such control lines themselves. Incoming messages without "CHRS" control lines should be considered as being written in the area's default character set (normally ASCII, IBM codepage 437, LATIN-1 or similar). 3. Supported levels These levels are the one that are implemented in current software: Level 0 ------- This level is for messages containing of pure seven-bit ASCII only. Outgoing messages in pure ASCII need not be identified by a "CHRS" control line, but if they are, they should be indicated as "ASCII 1" (not "ASCII 0"). Level 1 ------- First level of internationalization, using seven bit character sets. Most of these are based on US ASCII, with minor internationalization variations. Level 2 ------- Second level of internationalization, using eight bit character sets. This level adds support for character sets that use "extended ASCII", i.e codes with the most significant bit set. The character sets in level two are all based on ASCII (the codes 0-127 coincides with ASCII). Level 3 and above ----------------- Level 3 or higher as specified by FSC-54 has no known current implementation, because of which this document are considered reserved for future use. Level 3 would be suited for a Unicode (ISO 10646) implementation. Because of limitations of currently available Fidonet packet standards, it is not at the time of writing viable to specify how this should work. 4. Supported character sets These character sets are defined at the moment: Identifier Character set ---------- ------------- Level 1 character sets (seven-bit) ASCII ISO 646-1 (US ASCII) DUTCH ISO 646 Dutch FINNISH ISO 646-10 (Swedish/Finnish) FRENCH ISO 646 French CANADIAN ISO 646 Canadian GERMAN ISO 646 German ITALIAN ISO 646 Italian NORWEIG ISO 646 Norweigian PORTU ISO 646 Protuguese SPANISH ISO 646 Spanish SWEDISH ISO 646-10 (Swedish/Finnish) SWISS ISO 646 Swiss UK ISO 646 UK Level 1 known depreciated character set identifiers ISO-10 ISO 646-10 (Swedish/Finnish) Level 2 character sets (eight-bit, ASCII based) CP437 IBM codepage 437 (Western European) CP850 IBM codepage 850 (Latin-1) CP865 IBM codepage 865 (Nordic) CP866 IBM codepage 866 (Russian) LATIN-1 ISO 8859-1 (Western European) LATIN-2 ISO 8859-2 (Eastern European) LATIN-5 ISO 8859-9 (Turkish) MAC MacIntosh character set Level 2 obsolete character set identifiers (see note) IBMPC IBM PC character set +7_FIDO IBM codepage 866 5. Obsolete indentifiers Since the "IBMPC" identifier, initially used to indicate IBM codepage 437, eventually evolved into identifying "any IBM codepage", there exists in some implementations an additional control line, "CODEPAGE", identifying the messages codepage: ^ACODEPAGE: xxx This use is depreciated in favour of the "CPxxx" identifiers defined above. If found in incoming messages, however, it should be used as an override of the "CHRS: IBMPC" identifier. The character set "+7_FIDO" is currently in use as an identifier for CP866. This use is depreciated, and "CP866" is recommended instead. Implementations should treat "+7_FIDO" as a synonym for "CP866". FSC-54 also defined control lines for style changes, "CHRC". These are not implemented in any software known to the author, and are depreciated. Levels 3 and 4, as defined by FSC-54, are also considered "obsolete" since there currently are no known implementations of them. All levels not documented here are considered reserved for future use. 6. Notes The character set identifier applies to all parts of the message, including the header information, and the tear and Origin lines. FSC-54 documents file formats for mapping files that could be used for storing character translation data. These are not documented here. A. Author contact data Peter Karlsson Fidonet: 2:206/221 E-mail: pk@abc.se B. Acknowledgements Duncan McNutt (original FSC-0054 author). C. History Rev.1, 19990904: Based on FSC-0054, by Peter Karlsso