kodirovka sooschenij.
ZXNet echo conference «zxnet.pc»
From Kirill Frolov → To All 31 December 2003
Press RESET immediately, All!
I wrote a letter to ZX.SPECTRUM, went to this conference and immediately saw a couple
inappropriate letters. I decided to "forward".
In short: “+7 FIDO”, “IBMPC” and similar heresies do not have encodings
right to exist. You must use an encoding that is exactly
there are Russian letters. For example, CP866. ASCII does not contain Russian letters.
There is no such thing as high-ascii, ascii is by definition a 7-bit code.
Newsgroups: fido7.zx.spectrum
From: Kirill Frolov
X-Comment-To: All
Subject: Speaking Russian -- ZX Spectrum
Date: Wed Dec 31 09:38:24 MSK 2003
Press RESET immediately, All!
On Tue, 30 Dec 03 17:39:40 +0300, Andrei Denisenko wrote:
AD> A long time ago, in the 20s of the 20th century, Slavik Tretiak was building a machine gun on
AD> about: Р ї╜∙Г∙╛ ZX Spectrum Pojaluista po-russki.
This is what non-compliance with certain technical requirements leads to.
Let me quote both together with official information, on which
and will focus your attention. The letter is addressed to ALL ZX.SPECTRUM SUBSCRIBERS.
Please be careful and draw your own conclusions, check
settings for your programs.
So, the original letter (headers only):
Newsgroups: fido7.zx.spectrum
From: Slavik Tretiak
X-Comment-To: Andrei Denisenko
Subject: Development of ZX SpectrumDate: Tue, 30 Dec 03 11:41:02 +0300
Message-ID: <1072806062@p33.f2.n451.z2.FidoNet.ftn>
X-FTN-MSGID: 2:451/2.33@FidoNet 3ff1b8ae
X-FTN-CHRS: LATIN-1 2
X-FTN-Tearline: Spot 1.3b #118
This letter WAS NOT read by Andrey Denisenko because
the software he used showed it in ISO-8859-1 encoding, while
how the message text itself was encoded in CP866 encoding.
The program did *absolutely correctly* and in accordance with the relevant
standards (FTS) and recommendations (FTSC) of the Fidonet network. It's all about
in the service information line ("kludge") "X-FTN-CHRS: LATIN-1 2",
informing about the encoding used in the letter.
(Note: fidonet users may see this string as "kludge"
"@CHRS: LATIN-1 2", where the '@' character is a non-displayable character with code 01)
This string "kludge" contains INCORRECT TECHNICAL INFORMATION, therefore
that in reality the message is encoded in a different encoding.
I would like to note that this letter could not contain messages
in Russian, since the LATIN-1 encoding does not contain the necessary
characters.
Now let's look at Andrei Denisenko's letter, which was correct
must have been read by all teleconference subscribers (only
headers):
Newsgroups: fido7.zx.spectrum
From: Andrei Denisenko
X-Comment-To: Slavik Tretiak
Subject: Рї╜∙Г∙╛ ZX Spectrum
Date: Tue, 30 Dec 03 17:39:40 +0300Message-ID: <4215340360@p11.f7.n5064.z2.ftn>
X-FTN-MSGID: 2:5064/7.11 fb40fd48
What is visible here? The distorted theme immediately catches your eye
message, the original was “Development of the ZX Spectrum”. This is explained
the fact that Andrey's software is guided by something that does not correspond to reality
technical information incorrectly recoded the text. In addition,
it is clear that the CHRS “kludge” is missing. That is, it is unambiguous to determine
message encoding is not possible. Such messages have every right
to exist in the Fidonet network and meet all the necessary requirements
standards. However, programs processing such messages must
make any assumptions about the text encoding system. Usually
it is assumed that the one generally accepted for a given network or region is used
or network zone encoding. This works well for messages written
in the language accepted in this part of the network, for locally distributed
letters. But when sending mail (NETMAIL) to another zone, in international
teleconferences, or in teleconferences where it is customary to use
multiple languages, this may cause a problem -- assumptions
letters about encoding may be false.
So there is a problem: an undefined message encoding. And there is for
its solution provided by Fidonet standards - use
lines of technical information of the message header ("kludge")
"CHRS". This "kludge" must contain the official nametext coding systems. For the Russian part of the Fidonet network
It is accepted to use CP866 encoding. Ukraine, Belarus and
other republics of the former USSR may use other encodings,
only partially identical to CP866, for example CP1125. In European
parts of the 2nd zone of the Fidonet network and in the 1st zone (USA) it is common to use
CP855 or ISO-8859-1 encodings (may be referred to as LATIN-1).
Theoretically, a letter can be encoded in any of the broad
known (not necessarily standard - standard are
only encodings with the ISO prefix, for example, for Russian
language, only the ISO-8859-5 encoding is defined - this is the standard number),
including, the message can be encoded in unicode using
UTF-8 encoding system (use of multibyte encoding systems
unicode, like USC-16 is not possible in Fidonet). And, again theoretically,
messages encoded in any of the known encodings and containing
correct "kludge" CHRS, can be displayed correctly on the screen
recipient. So when using unicode it is possible to use
The letter contains several languages, for example Japanese and Russian. But that's all in
theory, in practice this is hampered by a number of restrictions imposed
outdated software that does not support work with
various encodings. This is why it is customary to encode
messages in some encoding generally accepted for the network, so that all subscribersnetworks, including those using outdated software, could easily
read your message.
So, what conclusions can be drawn from what has been said: “kludge” CHRS
MUST CONTAIN RELIABLE INFORMATION, or be missing
at all. Otherwise, your letter may be misinterpreted.
If this “kludge” is missing, the letter MUST BE
ENCODED IN A COMMON ENCODING FOR YOUR NETWORK. If "kludge"
present IT IS NOT ADVISABLE TO USE CODINGS OTHER THAN
COMMONLY ACCEPTED IN YOUR NETWORK, because not all subscribers are ready to accept
letters like this. So, in particular, a large network node 2:5020/400, through
which internet users have the ability to access
teleconferencing of the Fidonet network, uses software
does not support text recoding according to the CHRS "kludge",
and always expects that your letters will be encoded in generally accepted
for network encoding. If you encode your letters in excellent
from the generally accepted encoding, users of internet news servers
will not be able to read your letters.
REMEMBER: IF THE CHRS KLADGE CONTAINS INCORRECT INFORMATION, YOUR
THE LETTER MAY BE INCORRECTLY DECODED BY THE MESSAGE RECIPIENT,
SO IT IS DISTORTED AT INTERMEDIATE NODES_ WHICH CLEARLY FOLLOW
STANDARDS.
Recommendations from the author of this letter for setting up your software canbe like this: make sure that the CHRS "kludge" is present and always
contained correct information. So for Russia and most of the former USSR
this will be the generally accepted CP866 encoding. Try without extreme pressure
the need not to write letters in an encoding different from the generally accepted one.
If you cannot ensure that the CHRS "kludge" is correct, as in the case
the first of the messages discussed in this letter, MANDATORY
make sure that this “kludge” is not present at all in your
messages (even if the software you use does not allow it, "kludge"
can be cut with third-party programs), and always encode your letters
into a generally accepted encoding. If you need to put text in language,
which cannot be represented by a generally accepted encoding, encode
it in UUE or Base-64.
Many subscribers may have questions about setting up their
programs. First of all, carefully study the documentation, if it
no, then think about changing the software. So the latest version of the editor
GoldED (1.1.4.7 or later) fully supports various
encodings, but does not work with unicode. You can ask questions at
Fidonet network teleconferences "SU.CHAINIK" or "RU.UNIX.FTN" (for
using unix-like systems).
As an attachment to the letter, I provide document FSP1013, which describes
issues of internationalization in the Fidonet network. It might not be the last
document and contains inaccuracies.**********************************************************************
FTSC FIDONET TECHNICAL STANDARDS COMMITTEE
**********************************************************************
Publication: FSP-1013
Revision: 1
Title: Character set definition in Fidonet messages
Author: Peter Karlsson
Revision Date: 04 September 1999
Expiry Date: 04 September 2199
Contents:
1. Introduction
2. Format of the identifier
3. Supported levels
4. Supported character sets
5. Obsolete identifiers
6. Notes
Status of this document
This document suggests a proposed protocol for the FidoNet(r)
community, and requests discussion and suggestions for improvements.
Distribution of this document is unlimited.
Abstract
This document defines the identifiers that are used to indicate the
character set used within messages that are distributed in Fidonet.
There has been many attempts on defining a common standard on what
character set is used in Fidonet distributed messages. The only one
that has gained widespread is the "CHRS" standard, described in
FSC-0054. This document tries to describe the current usage, as
well as standardizing the parts of it that was ambigously defined.
1. Introduction
As Fidonet is an international network, one has to consider that
not all people speak English. Many languages also have alphabets
that are either bigger than the standard English alphabet, or
completely different. To keep track of which character set is used
for a particular message, this documents describes a way to identify
this to the reader.
2. Format of identifier
The CHRS control line is formatted as follows:
^ACHRS:
Where is a character string of no more than eight (8)
characters identifying the character set to use, and level is a
numeric value describing what level of CHRS this message is written
in.
For backward compatibility, the use of "CHARSET" instead of "CHRS"
should be supported when reading messages.
Some current implementations do not add the required field.
Implemetations of this document should allow for this usage, but
never write such control lines themselves.
Incoming messages without "CHRS" control lines should be considered
as being written in the area's default character set (normally
ASCII, IBM codepage 437, LATIN-1 or similar).
3. Supported levels
These levels are the one that are implemented in current software:
Level 0
-------
This level is for messages containing of pure seven-bit ASCII only.
Outgoing messages in pure ASCII need not be identified by a "CHRS"
control line, but if they are, they should be indicated as
"ASCII 1" (not "ASCII 0").
Level 1
-------
First level of internationalization, using seven bit character sets.
Most of these are based on US ASCII, with minor internationalization
variations.
Level 2
-------
Second level of internationalization, using eight bit character
sets.
This level adds support for character sets that use "extended
ASCII", i.e codes with the most significant bit set. The character
sets in level two are all based on ASCII (the codes 0-127 coincides
with ASCII).
Level 3 and above
-----------------
Level 3 or higher as specified by FSC-54 has no known current
implementation, because of which this document are considered
reserved for future use.
Level 3 would be suited for a Unicode (ISO 10646) implementation.
Because of limitations of currently available Fidonet packet
standards, it is not at the time of writing viable to specify how
this should work.
4. Supported character sets
These character sets are defined at the moment:
Identifier Character set
---------- -------------
Level 1 character sets (seven-bit)
ASCII ISO 646-1 (US ASCII)
DUTCH ISO 646 Dutch
FINNISH ISO 646-10 (Swedish/Finnish)
FRENCH ISO 646 French
CANADIAN ISO 646 Canadian
GERMAN ISO 646 German
ITALIAN ISO 646 Italian
NORWEIG ISO 646 Norweigian
PORTU ISO 646 Protuguese
SPANISH ISO 646 Spanish
SWEDISH ISO 646-10 (Swedish/Finnish)
SWISS ISO 646 Swiss
UK ISO 646 UK
Level 1 known depreciated character set identifiers
ISO-10 ISO 646-10 (Swedish/Finnish)
Level 2 character sets (eight-bit, ASCII based)
CP437 IBM codepage 437 (Western European)
CP850 IBM codepage 850 (Latin-1)
CP865 IBM codepage 865 (Nordic)
CP866 IBM codepage 866 (Russian)
LATIN-1 ISO 8859-1 (Western European)
LATIN-2 ISO 8859-2 (Eastern European)
LATIN-5 ISO 8859-9 (Turkish)
MAC MacIntosh character set
Level 2 obsolete character set identifiers (see note)
IBMPC IBM PC character set
+7_FIDO IBM codepage 866
5. Obsolete indentifiers
Since the "IBMPC" identifier, initially used to indicate IBM
codepage 437, eventually evolved into identifying "any IBM
codepage", there exists in some implementations an additional
control line, "CODEPAGE", identifying the messages codepage:
^ACODEPAGE: xxx
This use is depreciated in favour of the "CPxxx" identifiers
defined above. If found in incoming messages, however, it should be
used as an override of the "CHRS: IBMPC" identifier.
The character set "+7_FIDO" is currently in use as an identifier for
CP866. This use is depreciated, and "CP866" is recommended instead.
Implementations should treat "+7_FIDO" as a synonym for "CP866".
FSC-54 also defined control lines for style changes, "CHRC". These
are not implemented in any software known to the author, and are
depreciated.
Levels 3 and 4, as defined by FSC-54, are also considered "obsolete"
since there currently are no known implementations of them. All
levels not documented here are considered reserved for future use.
6. Notes
The character set identifier applies to all parts of the message,
including the header information, and the tear and Origin lines.
FSC-54 documents file formats for mapping files that could be used
for storing character translation data. These are not documented
here.
A. Author contact data
Peter Karlsson
Fidonet: 2:206/221
E-mail: pk@abc.se
B. Acknowledgements
Duncan McNutt (original FSC-0054 author).
C. History
Rev.1, 19990904: Based on FSC-0054, by Peter Karlsso